DEV Community

#cuda

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 in Pure JAX: What Changes Between Turing and Ada, and What Doesn't

Gemma 4 in Pure JAX: What Changes Between Turing and Ada, and What Doesn't

Comments
8 min read
g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput

g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput

Comments
6 min read
The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

Comments
11 min read
Accelerating Physical AI Workloads with NVIDIA CUDA and TensorRT

Accelerating Physical AI Workloads with NVIDIA CUDA and TensorRT

Comments
2 min read
Finding a Random Island with Geometry and CUDA

Finding a Random Island with Geometry and CUDA

Comments
4 min read
Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Comments
10 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Comments
9 min read
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU

Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU

1
Comments 1
11 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

1
Comments
2 min read
Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

1
Comments 1
3 min read
The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

Comments
3 min read
I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware

I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware

2
Comments
4 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

13
Comments
9 min read
Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization

Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization

Comments
2 min read
Serving Gemma4 with Rust on vLLM 🦀

Serving Gemma4 with Rust on vLLM 🦀

10
Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.