Opening the archive
Everything published here so far. Filter by topic, or sort by what gets read most.
Full parameter training requires clusters of H100s. Parameter-Efficient Fine-Tuning (PEFT) with low-rank matrix decomposition brings specialized adaptation to a single GPU.
Running self-hosted LLMs used to mean stitching vLLM, Triton, and CUDA drivers by hand. NVIDIA NIM packages TensorRT-LLM in standardized, production-ready OCI containers.
A web-based prototyping environment with foundation models, knowledge bases, agents, and guardrails. Public preview, two regions.
Pinecone Serverless for sub-50ms vector retrieval, Google's Gemini for grounded synthesis. A complete end-to-end Python implementation with metadata filtering.
Massive context windows, native multimodality, and structured schema outputs: how developers can build production systems on Google AI Studio.
They all install packages. Where they diverge is the shape of node_modules, and that difference has consequences for disk space and for correctness.
735 views
728 views
683 views
356 views