Opening the archive
Everything published here so far. Filter by topic, or sort by what gets read most.
RAG searches for information on every turn. CAG pre-computes the KV cache. Where Time-to-First-Token and conversational latency matter, caching flips the architecture.
Language Processing Units are purpose-built for language models rather than general compute. The specialization shows up in where the weights live.
735 views
728 views
683 views
356 views