Opening the archive
Everything published here so far. Filter by topic, or sort by what gets read most.
RAG searches for information on every turn. CAG pre-computes the KV cache. Where Time-to-First-Token and conversational latency matter, caching flips the architecture.
Tokens, vectors, attention, and next-token prediction. What is actually happening underneath something that feels like a conversation.
Language Processing Units are purpose-built for language models rather than general compute. The specialization shows up in where the weights live.
ChatGPT did not cause the wildfires. The conversation that formed around the claim is still worth having.
There are not enough GPUs to go around, so the company behind ChatGPT is looking at making the hardware itself.
Gemma is capable and open source. Hugging Face is what makes getting to it straightforward.
735 views
728 views
683 views
356 views