Opening the archive
Everything published here so far. Filter by topic, or sort by what gets read most.
Gemma is capable and open source. Hugging Face is what makes getting to it straightforward.
Full parameter training requires clusters of H100s. Parameter-Efficient Fine-Tuning (PEFT) with low-rank matrix decomposition brings specialized adaptation to a single GPU.
Running self-hosted LLMs used to mean stitching vLLM, Triton, and CUDA drivers by hand. NVIDIA NIM packages TensorRT-LLM in standardized, production-ready OCI containers.
A web-based prototyping environment with foundation models, knowledge bases, agents, and guardrails. Public preview, two regions.
Pinecone Serverless for sub-50ms vector retrieval, Google's Gemini for grounded synthesis. A complete end-to-end Python implementation with metadata filtering.
Massive context windows, native multimodality, and structured schema outputs: how developers can build production systems on Google AI Studio.
735 views
728 views
683 views
356 views