News
Latest
Top
Search
Submit
Login
Search
▲
271
CUDA Ontology
(jamesakl.com)
by gugagore |
view
|
40 comments
▲
132
CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
(github.com)
by dzign |
view
|
15 comments
▲
88
Why CUDA translation wont unlock AMD
(eliovp.com)
by JonChesterfield |
view
|
81 comments
▲
22
Parrot – A C++ library for fused array operations using CUDA/Thrust
(nvlabs.github.io)
by operator-name |
view
|
2 comments
▲
17
Apex GPU: Run CUDA Apps on AMD GPUs Without Recompilation
(github.com)
by ArchitectAI |
view
|
6 comments
▲
11
Show HN: CUDA, Shmuda: Fold Proteins on a MacBook
(latentspacecraft.com)
by geoffitect |
view
|
3 comments
▲
6
Nvidia CUDA Tile
(developer.nvidia.com)
by elkguy |
view
|
0 comments
▲
5
Show HN: Clangd for CUDA Device Code
(docs.scale-lang.com)
by JonChesterfield |
view
|
0 comments
▲
5
Nvidia's B200: Keeping the CUDA Juggernaut Rolling Ft Verda, Formerly DataCrunch
(chipsandcheese.com)
by rbanffy |
view
|
0 comments
▲
5
Nvidia cuTile: Python DSL and a new IR for tile-based CUDA kernels
(github.com)
by ashvardanian |
view
|
0 comments
▲
4
NVIDIA CUDA Tile programming model
(developer.nvidia.com)
by tanelpoder |
view
|
0 comments
▲
4
Show HN: modal-cuda – CLI to run CUDA .cu programs on Modal GPUs
(github.com)
by Sai_Praneeth |
view
|
0 comments
▲
3
AI-Written CUDA Kernels Outperforms Nvidia's Best Matmul Library
(rohan-paul.com)
by dzign |
view
|
0 comments
▲
3
An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
(arxiv.org)
by PaulHoule |
view
|
0 comments
▲
3
Understanding the CUDA Compiler and PTX with a Top-K Kernel
(blog.alpindale.net)
by mfiguiere |
view
|
0 comments
▲
2
AMD calls CUDA a 'non-event'
(pcgamer.com)
by neilfrndes |
view
|
0 comments
▲
2
Performant C/CUDA inference engine for Qwen 3.6 35B on RTX 5090 / Blackwell
(github.com)
by ambuds |
view
|
0 comments
▲
2
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(github.com)
by vforno |
view
|
0 comments
▲
2
Show HN: CUDA Profiler for Production Inference
(github.com)
by npgraph |
view
|
0 comments
▲
2
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
(arxiv.org)
by matt_d |
view
|
0 comments
▲
2
CUDA Released in Basic
(developer.nvidia.com)
by apples2apples |
view
|
0 comments
▲
2
Nvidia CUDA Tile
(developer.nvidia.com)
by apples2apples |
view
|
0 comments
▲
2
CUDA Tile
(techpowerup.com)
by dagmx |
view
|
0 comments
▲
2
Show HN: Free GPUs in your terminal for learning CUDA
(github.com)
by RohanAdwankar |
view
|
0 comments
▲
2
CUDA-Q Back Ends: Quantum Hardware (QPU)
(nvidia.github.io)
by westurner |
view
|
3 comments
▲
1
Nvprobe – Open-source, zero-setup CLI for CUDA benchmarks
(github.com)
by SergioZ3R0 |
view
|
0 comments
▲
1
Bw24 – from scratch rust+CUDA inference, every kernel tuned for sm_120a
(github.com)
by anotherCodder |
view
|
0 comments
▲
1
Running CUDA on Apple GPUs
(twitter.com)
by abhinavsns |
view
|
0 comments
▲
1
Eliminating Conda/CUDA dependency hell in computational biology pipelines
(github.com)
by LionTurtle13 |
view
|
0 comments
▲
1
CUDA-accelerated program to search Minecraft seeds for Mooshrom Island biomes
(github.com)
by Tmpod |
view
|
0 comments
▲
1
Show HN: Trellis2.c – Local 3D generation with Vulkan and CUDA
(github.com)
by wimaxs |
view
|
0 comments
▲
1
Fusing a 27B ternary LLM's whole decode step into one CUDA kernel
(twitter.com)
by Jr23_xd |
view
|
0 comments
▲
1
Alternative(s) to run CUDA on non-Nvidia hardware
(hpcwire.com)
by alok-g |
view
|
0 comments
▲
1
Optimizing CUDA Like a Human: Micro-Profiling Tools as Expert Surrogates For
(hgpu.org)
by ibobev |
view
|
0 comments
▲
1
What If Java Apps Could Access CUDA Ecosystem Gracefully
(tornadovm.org)
by mikepapadim |
view
|
0 comments
▲
1
1970 Plymouth Hemi 'CUDA
(knuckledustchronicles.com)
by frobinson47 |
view
|
0 comments
▲
1
Show HN: A 100% branchless, CUDA-native AI guardrail kernel written in C++20
(github.com)
by PJHkorea |
view
|
0 comments
▲
1
Show HN: Real-time n-body tree code in CUDA
(github.com)
by lechebs |
view
|
0 comments
▲
1
Reverse-engineering Nvidia's CUDA-checkpoint for faster cold starts
(blog.doubleword.ai)
by ilreb |
view
|
0 comments
▲
1
Show HN: A free, GPU-accelerated Texas Hold'em GTO solver in C++/CUDA
(bupticybee.github.io)
by bupticybee |
view
|
0 comments
▲
1
Zluda (CUDA-compatible runtime for AMD) loses funding again, gains 32bit compat
(vosen.github.io)
by indrora |
view
|
0 comments
▲
1
Zluda 6 release (run unmodified CUDA applications on non-Nvidia GPUs)
(vosen.github.io)
by Tiberium |
view
|
0 comments
▲
1
What happens when you run a CUDA kernel?
(fergusfinn.com)
by mezark |
view
|
0 comments
▲
1
Optimizing a CUDA FSST decompression kernel
(polarsignals.com)
by asubiotto |
view
|
0 comments
▲
1
CUDA Profiler for Production Inference
(graphsignal.com)
by npgraph |
view
|
0 comments
▲
1
Running a 35B MoE model on a 2017 AMD RX 580 8GB via Vulkan (no ROCm/CUDA)
(github.com)
by aivisionslab |
view
|
0 comments
▲
1
Fast Great-Circle Distance Calculation in CUDA C++
(developer.nvidia.com)
by Alien1Being |
view
|
0 comments
▲
1
Running local AI on AMD RX 580 (2017 GPU) using Vulkan – no CUDA, no ROCm
(setup-ia-local-rx580-vulkan.web.app)
by aivisionslab |
view
|
0 comments
▲
1
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(github.com)
by vforno |
view
|
0 comments
▲
1
Show HN: FlashQwen – A from-scratch CUDA inference engine for Qwen3
(github.com)
by langtang1996 |
view
|
0 comments