DEV Community

#inference

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Training vs Inference: Why Building Costs Millions and Asking Costs Cents

Training vs Inference: Why Building Costs Millions and Asking Costs Cents

Comments
9 min read
I Got 28 TPS Out of Free Kaggle GPUs. Here's What It Took.

I Got 28 TPS Out of Free Kaggle GPUs. Here's What It Took.

Comments
5 min read
Speculative Decoding and MTP: Why Guessing Is Free

Speculative Decoding and MTP: Why Guessing Is Free

Comments
6 min read
Build and Evaluate an AI Error Explainer with DigitalOcean Inference

Build and Evaluate an AI Error Explainer with DigitalOcean Inference

Comments
12 min read
what a turn actually costs me

what a turn actually costs me

Comments
2 min read
AMD's Move on Weight Storage: The Taalas Bet

AMD's Move on Weight Storage: The Taalas Bet

Comments
2 min read
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

1
Comments 1
4 min read
AMD Bets Weight Storage Is the Real Bottleneck

AMD Bets Weight Storage Is the Real Bottleneck

Comments
2 min read
vLLM reinvented the operating system, and nobody told you

vLLM reinvented the operating system, and nobody told you

1
Comments
12 min read
KV Cache by hand

KV Cache by hand

1
Comments
11 min read
Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency

Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency

Comments
2 min read
LLM Inference Optimization: From Quantization to Speculative Decoding

LLM Inference Optimization: From Quantization to Speculative Decoding

1
Comments 1
2 min read
local-llm: A Field Report on Running SOTA Models on Your Own Hardware

local-llm: A Field Report on Running SOTA Models on Your Own Hardware

1
Comments 1
3 min read
The KV cache, why LLM inference is memory-bound, not compute-bound

The KV cache, why LLM inference is memory-bound, not compute-bound

Comments
4 min read
Etched hits $5B and $1B in orders: why inference chips matter

Etched hits $5B and $1B in orders: why inference chips matter

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.