Fig. I — The almanac index

Tags

Every field note, cross-referenced by topic. For the archive's primary divisions see the categories, and for the runs meant to be read in order, the series.

All tags

explainer (19)quantization (8)inference (6)training data (5)lora (5)fine-tuning (5)reinforcement learning (5)open models (4)evaluation (4)large language models (4)vector space (3)embeddings (3)scaling laws (3)transformers (3)gpus (2)machine learning (2)kimi k3 (2)mixture of experts (2)ai safety (2)building blocks (2)attention (2)architecture (2)distillation (2)synthetic data (2)data curation (2)reasoning (1)decoding (1)stochastic resonance (1)parallel computing (1)tensor cores (1)hardware (1)artificial intelligence (1)ani agi asi (1)sandbox escape (1)security (1)hugging face (1)reward hacking (1)specification gaming (1)series closing (1)training (1)forward pass (1)tokenization (1)byte-pair encoding (1)vocabulary (1)tensors (1)shapes (1)dimensions (1)semantic similarity (1)vectors (1)linear algebra (1)representation (1)weights (1)parameters (1)neural networks (1)series introduction (1)agi (1)ai forecasting (1)open weights (1)ai policy (1)bias (1)fairness (1)ai ethics (1)automation (1)the labour market (1)ai and work (1)ai agents (1)tool use (1)autonomy (1)prompt engineering (1)generative ai (1)diffusion models (1)image generation (1)chatgpt (1)next-token prediction (1)moonshot ai (1)pruning (1)optimization (1)post-mortem (1)hallucination (1)vision models (1)ocr (1)ai systems (1)self-play (1)onnx (1)modal (1)training pipelines (1)rlhf (1)dpo (1)alignment (1)reasoning models (1)ppo (1)policy gradients (1)trpo (1)dqn (1)q-learning (1)experience replay (1)reward functions (1)markov decision process (1)qat (1)bitnet (1)1-bit models (1)gguf (1)local models (1)llama.cpp (1)gptq (1)awq (1)outliers (1)int4 (1)numerical precision (1)model serving (1)adapters (1)hyperparameters (1)rank (1)qlora (1)nf4 (1)low-rank adaptation (1)parameter-efficient tuning (1)instruction tuning (1)model collapse (1)data quality (1)data pipelines (1)deduplication (1)common crawl (1)semantic search (1)cosine similarity (1)chinchilla (1)compute budgets (1)model size (1)self-attention (1)gradient descent (1)backpropagation (1)optimizers (1)loss functions (1)