News
Latest
Top
Search
Submit
Login
Search
▲
271
CUDA Ontology
(jamesakl.com)
by gugagore |
view
|
40 comments
▲
132
CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
(github.com)
by dzign |
view
|
15 comments
▲
88
Why CUDA translation wont unlock AMD
(eliovp.com)
by JonChesterfield |
view
|
81 comments
▲
22
Parrot – A C++ library for fused array operations using CUDA/Thrust
(nvlabs.github.io)
by operator-name |
view
|
2 comments
▲
17
Apex GPU: Run CUDA Apps on AMD GPUs Without Recompilation
(github.com)
by ArchitectAI |
view
|
6 comments
▲
11
Show HN: CUDA, Shmuda: Fold Proteins on a MacBook
(latentspacecraft.com)
by geoffitect |
view
|
3 comments
▲
6
Nvidia CUDA Tile
(developer.nvidia.com)
by elkguy |
view
|
0 comments
▲
5
Show HN: Clangd for CUDA Device Code
(docs.scale-lang.com)
by JonChesterfield |
view
|
0 comments
▲
5
Nvidia's B200: Keeping the CUDA Juggernaut Rolling Ft Verda, Formerly DataCrunch
(chipsandcheese.com)
by rbanffy |
view
|
0 comments
▲
5
Nvidia cuTile: Python DSL and a new IR for tile-based CUDA kernels
(github.com)
by ashvardanian |
view
|
0 comments
▲
4
NVIDIA CUDA Tile programming model
(developer.nvidia.com)
by tanelpoder |
view
|
0 comments
▲
4
Show HN: modal-cuda – CLI to run CUDA .cu programs on Modal GPUs
(github.com)
by Sai_Praneeth |
view
|
0 comments
▲
3
AI-Written CUDA Kernels Outperforms Nvidia's Best Matmul Library
(rohan-paul.com)
by dzign |
view
|
0 comments
▲
3
An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
(arxiv.org)
by PaulHoule |
view
|
0 comments
▲
3
Understanding the CUDA Compiler and PTX with a Top-K Kernel
(blog.alpindale.net)
by mfiguiere |
view
|
0 comments
▲
2
AMD calls CUDA a 'non-event'
(pcgamer.com)
by neilfrndes |
view
|
0 comments
▲
2
Performant C/CUDA inference engine for Qwen 3.6 35B on RTX 5090 / Blackwell
(github.com)
by ambuds |
view
|
0 comments
▲
2
Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(github.com)
by vforno |
view
|
0 comments
▲
2
Show HN: CUDA Profiler for Production Inference
(github.com)
by npgraph |
view
|
0 comments
▲
2
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
(arxiv.org)
by matt_d |
view
|
0 comments
▲
2
CUDA Released in Basic
(developer.nvidia.com)
by apples2apples |
view
|
0 comments
▲
2
Nvidia CUDA Tile
(developer.nvidia.com)
by apples2apples |
view
|
0 comments
▲
2
CUDA Tile
(techpowerup.com)
by dagmx |
view
|
0 comments
▲
2
Show HN: Free GPUs in your terminal for learning CUDA
(github.com)
by RohanAdwankar |
view
|
0 comments
▲
2
CUDA-Q Back Ends: Quantum Hardware (QPU)
(nvidia.github.io)
by westurner |
view
|
3 comments
▲
1
Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
(developer.nvidia.com)
by ubj |
view
|
0 comments
▲
1
Show HN: CUDA/graphics in QEMU-KVM VMs without passing the Nvidia card to them
(github.com)
by reindertpelsma |
view
|
0 comments
▲
1
CUDA Rust: Two Tracks for Writing GPU Kernels
(developer.nvidia.com)
by vertigoruntime |
view
|
0 comments
▲
1
I have made an OS in CUDA
(twitter.com)
by porridgeraisin |
view
|
0 comments
▲
1
CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
(developer.nvidia.com)
by elashri |
view
|
0 comments
▲
1
Show HN: CuMetal – Run CUDA Programs on Apple Silicon via Metal
(github.com)
by lulzx |
view
|
0 comments
▲
1
Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam
(chipsandcheese.com)
by rbanffy |
view
|
0 comments
▲
1
Writing Fast Attention Backward Kernel for 5090 in CUDA
(hoanh.space)
by hoanh |
view
|
0 comments
▲
1
Cudametal – vibecoded Python translator from .cu to mpl
(github.com)
by mkunc |
view
|
0 comments
▲
1
Cudametal vibe coded Python tool
by mkunc |
view
|
0 comments
▲
1
Fork of antirez h3.c highly optimized for CUDA and DGX Spark (~15.5x speedup)
(github.com)
by binyu |
view
|
0 comments
▲
1
Testing CUDA kernel execution with GPU correctness harness
(github.com)
by laxmena |
view
|
0 comments
▲
1
Geolocating a random island using geometry and CUDA programming
(yassa9.github.io)
by yassa9 |
view
|
0 comments
▲
1
Stanford's deterministic CUDA kernel verifier
(2026.splashcon.org)
by ggboimoney |
view
|
0 comments
▲
1
Java at the Metal: CUDA Graphs, Tensor Cores, cuBLAS/CuDNN/CuFFT with TornadoVM
(tornadovm.org)
by pjmlp |
view
|
0 comments
▲
1
CUDA Shared Memory Swizzling
(leimao.github.io)
by jxmorris12 |
view
|
0 comments
▲
1
Project Kalos – Zero-copy C/CUDA sidecar for 0.46ms LLM memory recall
(github.com)
by mongoosereborn |
view
|
0 comments
▲
1
Kernel Fusion in Nvidia CUDA: Optimizing Memory Traffic and Launch Overhead
(developer.nvidia.com)
by fzimmermann89 |
view
|
0 comments
▲
1
Show HN: G6k-rs – lattice reduction framework for rust CPU,Metal,CUDA
(github.com)
by _alphageek |
view
|
0 comments
▲
1
Nvidia's CUDA Faces New Threats from AI Coding Agents
(businessinsider.com)
by chrsw |
view
|
0 comments
▲
1
Foundational ternary-model inference and training – CUDA, CPU, BitNet/TQ
(github.com)
by jacquesm |
view
|
0 comments
▲
1
Can AMD Break the CUDA Moat? AMD Advancing AI 2026
(newsletter.semianalysis.com)
by rbanffy |
view
|
0 comments
▲
1
HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA
(github.com)
by XXXGhost |
view
|
0 comments
▲
1
From CUDA to MLX: How K-Search Brings Kernel Expertise to Apple Silicon
(bair.berkeley.edu)
by williamjinq |
view
|
0 comments
▲
1
Can AMD Break the CUDA Moat? AMD Advancing AI 2026
(newsletter.semianalysis.com)
by gmays |
view
|
0 comments