Skip to content
Tag

Everything tagged #cag

1 article, newest first.

InferenceFeb 12, 2025 · 3 min read

CAG Over RAG, When Speed Is the Constraint

RAG searches for information on every turn. CAG pre-computes the KV cache. Where Time-to-First-Token and conversational latency matter, caching flips the architecture.

003 min read
CAG Over RAG, When Speed Is the Constraint