AI Model Optimization

A Study Guide
Knowledge Distillation  •  NAS  •  Pruning  •  Quantization  •  Low-Rank Factorization  •  Weight Sharing  •  Production Pipelines

Contents

  1. Knowledge Distillation
  2. Neural Architecture Search (NAS)
  3. Pruning
  4. Quantization
  5. Low-Rank Factorization & LoRA
  6. Weight Sharing
  7. Real Production Pipelines
  8. Early Exit
  9. Combining Methods
  10. References
Part 1

Knowledge Distillation

What is it?

Knowledge distillation is a technique where a small model (the student) learns by imitating a large, accurate model (the teacher), rather than learning directly from raw labeled data alone. The goal is to transfer the teacher's "knowledge" into a compact model that is cheaper to run.

Figure 1.1 — Distillation overview: teacher guides student via soft labels
Input imagee.g. a photo of a cat Teacher modelLarge, slow, accuratee.g. 100 billion params Student modelSmall, fast, lightweighte.g. 1 billion params Soft labels (from teacher)cat: 85% dog: 12% fox: 3%Rich probability distribution Hard labels (ground truth)cat: 100% dog: 0% fox: 0%One correct answer only Distillation lossMix of soft + hard → trains student

Hard labels vs Soft labels

Hard labels are human-annotated ground truth: one class is 100%, all others are 0%. They carry no information about similarity between classes.

Soft labels are the probability distributions output by the teacher model (e.g. cat 85%, fox 12%, dog 3%). They capture which classes resemble each other — information invisible in hard labels. Hinton called this dark knowledge.

Figure 1.2 — Hard vs soft labels side by side
Hard labels "One winner, everything else is zero" Golden retriever100% Labrador0% Wolf0% Fox0% Comes from a human annotator. Binary. No similarity information. Soft labels "Confident, but aware of similarity" Golden retriever78% Labrador15% Wolf4% Fox2% Comes from the teacher model. Encodes similarity between classes. Key insight: soft labels reveal which classes are similar to each other — hard labels hide this.

The distillation loss formula

Total loss = α × T² × KL(teacher ∥ student) + (1 − α) × CrossEntropy(student, true label)
  • Soft loss (KL divergence): measures how different the student's distribution is from the teacher's. Both outputs are passed through temperature T.
  • Hard loss (cross-entropy): standard training loss vs ground-truth label. Keeps the student grounded.
  • Alpha (α): mixing weight, typically 0.7–0.9. Higher = trust the teacher more.
  • T² scaling: compensates for gradient shrinkage caused by raising temperature.
Figure 1.3 — Distillation loss breakdown
Total loss = α × (soft loss) + (1 − α) × (hard loss) α is a mixing weight you choose, typically 0.7–0.9 Soft loss (KL divergence)Compare student vs teacher distributionsBoth softened by temperature T"How different are their beliefs?" Hard loss (cross-entropy)Compare student vs ground truthNo temperature used here"Did the student get it right?" × α × (1−α) Total distillation lossUpdates student weights via backpropagation Note: soft loss is also scaled by T² to compensate for gradient shrinkage
Part 2

Neural Architecture Search (NAS)

What is it?

NAS automates the design of neural network architectures. Instead of a human expert manually choosing layers, layer types, and connections, an algorithm searches for the best combination automatically.

Figure 2.1 — NAS three core components
① Search spaceAll possible networkdesigns to considerHow many layers? Which ops? ② Search strategyHow to explore thesearch space smartlyRL / Evolutionary / DARTS ③ EstimatorQuickly evaluate howgood each candidate isWithout full training explores score fed back Best architecture foundTrained fully and deployed How NAS relates to distillation NAS finds a good small architecture. Distillation then trains it efficiently. They are often used together: NAS designs the student, distillation teaches it.

Performance estimator tricks

Figure 2.2 — Three estimation tricks and their relative speed
The problem: fully training each candidate takes daysNAS must evaluate thousands of architectures → need fast approximations Early stoppingTrain for just 5 epochsinstead of 100. Seewhich looks mostpromising early on. Weight sharingTrain one giant"supernet" containingall candidates inside.Borrow weights. Predictor modelTrain a small ML modelto predict scoreswithout runningthe network at all. Relative speed vs full training Full training100× (baseline) Early stopping~10–20× faster Weight sharing~100–1000× faster Predictor model~10,000× faster Trade-off: speed vs accuracy of the estimate. NAS picks the right balance for the task.

Hardware-aware NAS: dominant schemes

MethodSearch costHardware-aware?Best for
DARTSLow (GPU-hours)Indirect (FLOPs)Fast prototyping
Once-for-AllHigh once, free afterYes — many devicesMulti-device deployment
ProxylessNASMediumYes — real latencySingle target chip
EvolutionaryVery high (GPU-days)Yes — real latencyProduction models
Part 3

Pruning

What is it?

Pruning removes parts of a trained neural network that contribute little to its output. 50–90% of weights in a typical trained network can be set to zero with negligible accuracy loss. The Lottery Ticket Hypothesis (2019) showed that inside every large network there is a small "winning sub-network" that does most of the real work.

Figure 3.1 — Before and after pruning: dense vs sparse network
Before pruning After pruning prune ✕ Dense: many weak connections (gray) Sparse: only strong connections remain

Three criteria for deciding what to prune

  • Magnitude-based: remove weights closest to zero. Cheapest, no data required. Assumes small weight = small contribution.
  • Gradient-based: remove weights whose removal changes the loss the least (estimated via gradient × weight). More principled, used in Taylor expansion pruning.
  • Activation-based: run real data through; remove neurons that are rarely activated. Natural fit for structured pruning.

The standard workflow

Figure 3.2 — Pruning workflow with iterative loop
① Trainfull network ② Prunezero out weights ③ Fine-tunerecover accuracy ④ Deploysmaller, faster model repeat: iterative pruning (prune a little, fine-tune, repeat) Sweet spot: 50–70% pruned with <1–2% accuracy drop using iterative approach
Part 4

Quantization

What is it?

Quantization reduces numerical precision of weights from 32-bit floats to 8-bit or 4-bit integers. Neural networks are surprisingly tolerant of imprecision — rounding 0.37219184 to an integer like 47 produces nearly identical outputs.

Figure 4.1 — Bit widths compared
FP3232 bitse.g. 0.37219184 — 1× (baseline) FP1616 bitse.g. 0.3721 — 2× smaller INT88 bits47e.g. 47 (out of −128…127) — 4× smaller INT44 bits3e.g. 3 (out of 0…15) — 8× smaller float = scale × (integer − zero_point) | store scale + zero_point alongside each layer

LLM quantization formats

FormatFull nameKey insightBest for
GPTQGeneralized Post-Training QuantizationCompensate rounding error in remaining weights layer by layerNVIDIA GPU deployment
AWQActivation-aware Weight QuantizationProtect the 1% of salient weights; aggressively compress the restNVIDIA GPU, highest accuracy
GGUFUniversal model file for llama.cpp (successor to GGML)Self-contained file for any hardware — CPU, GPU, Apple SiliconLaptops, Ollama, llama.cpp
Part 5

Low-Rank Factorization & LoRA

What is low-rank factorization?

A weight matrix W (e.g. 1000×1000 = 1,000,000 parameters) can be approximated by two thin matrices A (1000×r) and B (r×1000). At rank 4, A+B store only 8,000 parameters — 125× fewer.

Figure 5.1 — Matrix decomposition and LoRA adapter
W 1000×1000 1,000,000 params ≈ A 1000×4 × B 4×1000 Parameter savings W alone: 1,000,000 params A+B (rank 4): 8,000 params 125× fewer parameters LoRA: freeze W, train only A and B W (frozen) never updated + ΔW = A × B tiny trainable adapter Output = (W + ΔW) × x merged at deploy time Memory Fine-tune 70B on 1 GPU Swappable adapters 1 base, many task LoRAs Tiny file size Full: 140 GB → LoRA: 50 MB

How the loss is calculated with frozen W

Figure 5.2 — Gradient flow: forward pass uses W, backward pass skips it
→ Forward pass ← Backward pass (gradients) input x W (frozen) requires_grad=False W·x A (trainable) rank bottleneck B (trainable) projects back up B·A·x + output yW·x + B·A·x Loss updates B ✓ updates A ✓ blocked ✗ W never updated ✗ "Blocked" = requires_grad=False on W. PyTorch computes gradients up to W then stops. W is never changed.
QLoRA
Base model stored in INT4 + LoRA adapters trained in full precision. Fine-tune a 70B model on a single consumer GPU.
Part 6

Weight Sharing

What is it?

Weight sharing makes multiple weights share the same stored value. Instead of storing each weight as its own 32-bit number, you assign every weight to a cluster and store one representative value per cluster in a codebook. Each weight is replaced by a small index pointing to its cluster.

Figure 6.1 — Weight sharing via k-means codebook
Before: 16 unique weights After: 4 shared values + index grid 0.91 -0.52 0.12 -0.89 0.08 0.88 -0.93 -0.48 -0.55 0.15 0.94 -0.91 0.89 -0.87 -0.50 0.11 k-means 0 1 2 3 2 0 3 1 1 2 0 3 0 3 1 2 Codebook (4 entries) 0: +0.91 1: -0.51 2: +0.12 3: -0.90 Before: 16 × 32 bits = 512 bits stored After: 16 × 2-bit index + 4 × 32-bit codebook = 160 bits → 3× smaller Key difference from quantization: codebook values are learned (non-uniform), not a fixed grid
Part 7

Real Production Pipelines

No technique is used alone — they stack

In production, compression techniques are always combined. Each addresses a different dimension: distillation shrinks architecture, pruning removes redundancy, quantization reduces precision, LoRA enables efficient customization.

Figure 7.1 — Three real production pipelines
Pipeline A — On-device mobile (Apple Neural Engine, Pixel) NASfind arch Distillationtrain student Pruningstructured QuantizeINT8/INT4 Deploy on-chipruns in <10ms on device Pipeline B — Custom LLM fine-tuning (Llama / Mistral on your task) Pretrainedbase LLM FP16 Quantizebase to INT4 QLoRAadd adapters Merge+GGUFexport Ollama Deployruns on laptop Pipeline C — Cloud LLM inference at scale (serving GPT-class models) Large modelFP32 trained Distillationsmaller model Pruningstructured INT8 quant+calibration TensorRT/vLLMoptimized serving Three universal rules Order mattersNAS/distillation first.Quantization always last. Recover between stepsFine-tune after pruningbefore quantizing. Measure at each stepTrack accuracy drop.Don't let it compound. Distillation: 10× smaller + Pruning: 2× smaller + Quantization: 4× smaller Combined: ~80× smaller than original, with only ~2–5% total accuracy loss

Why order matters

Architecture-level decisions first. NAS and distillation define the shape and size of the model. You want to establish this foundation before applying lossy operations like pruning or quantization.

Quantization last. It is the most aggressive single-step transformation. Applying it to a model that has already been well-optimized by pruning and fine-tuning minimizes the total accuracy loss.

Recover between steps. After pruning and before quantizing, always fine-tune to recover lost accuracy. Stacking two lossy steps without recovery compounds the damage significantly.

Part 8

Early Exit

What is it?

Early exit adds off-ramps at intermediate layers of a neural network. Instead of every input running through all layers, easy inputs exit after fewer layers while hard inputs continue to the end. The key insight: not all inputs are equally difficult, so why spend the same compute on all of them?

Figure 8.1 — Early exit network: easy inputs leave early, hard inputs go deep
Input batch some inputs are easy • some are hard • we don't know which yet Layer block 1 (first few layers) Exit 1: confidence ≥ threshold? e.g. top class probability ≥ 0.90 EXIT ✓ easy input not confident enough → Layer block 2 (middle layers) Exit 2: confidence ≥ threshold? e.g. top class probability ≥ 0.90 EXIT ✓ medium input Layer block 3 (full network — all remaining layers) Final output (always exits here) hardest inputs — full compute used Key benefit: average compute per input drops dramatically If 60% of inputs exit at layer 6 of 48, average depth ≈ 25 layers instead of 48 Training: loss computed at every exit simultaneously, summed together. All exit heads train jointly with the main network.

How the confidence check works

Each exit point has a small classifier head — usually just one or two layers — attached to the main network. After processing through a block of layers, this head outputs a probability distribution. If the top predicted class has probability above a threshold (e.g. 0.90), the network is confident enough to exit now. If the distribution is flat ("20% cat, 18% dog, 15% fox..."), confidence is low and the input continues deeper.

The threshold is a tunable dial: higher = fewer exits (more accurate), lower = more exits (faster). You choose it based on how much accuracy you're willing to trade for speed.

Where early exit shines

Figure 8.2 — Average compute used relative to full network, by workload difficulty
Average compute used relative to full network No early exit 100% compute (baseline) Easy workload ~30% compute (70% saved) Mixed workload ~55% compute (45% saved) Hard workload ~85% compute (15% saved) Unlike pruning/quantization — savings are dynamic and input-dependent Works best when workloads are naturally skewed toward easy inputs Main limitation: doesn't play well with batched GPU inference — different inputs exit at different layers

Ideal use cases

  • NLP classification: "This email is spam" is obvious after 3 layers. No need for 12 layers.
  • Real-time video: Most frames show a routine scene. Run the full model only when something changes.
  • LLM token generation: Common words ("the", "and") are predictable early. Exit sooner for those tokens.
  • Edge / CPU inference: Where inputs arrive one at a time and batch parallelism isn't the bottleneck.

The key distinction from all other techniques

Every other method — pruning, quantization, distillation — produces a fixed speedup regardless of input. Early exit is fundamentally different: its savings are dynamic. The harder the batch, the less you save. This makes it complementary to the other techniques rather than a replacement — you apply pruning and quantization first for a baseline speedup, then layer early exit on top for additional input-adaptive savings.

Part 9 — Capstone

Combining Methods in Production

Start from your constraint, not the technique

The biggest mistake practitioners make is picking a technique because it sounds interesting rather than because it matches their actual problem. The table below maps your situation to the right starting point.

Your situationPrimary techniqueStack with
Need to run a large model on a laptop / phoneQuantization (INT4) — GGUF / AWQPruning for further size
Want to fine-tune a large LLM on your own dataLoRA / QLoRA — adapter-based fine-tuningINT4 base + distillation
Building a model for a specific chip / deviceNAS (hardware-aware) — ProxylessNAS / OFADistillation + INT8
Large model works well, want a smaller oneKnowledge distillation — teacher → studentPruning + quantization
Model is too slow, accuracy is fineStructured pruning — remove whole neuronsQuantization after
Workload has many "easy" inputsEarly exit — dynamic compute savingAny other technique first
Memory is the bottleneck, not speedWeight sharing — codebook compressionLow-rank factorization

Case Study A — Custom LLM running on a laptop

This is the most common real-world scenario for practitioners today. You don't train from scratch — that costs millions. Instead:

Figure 9.1 — Open-source LLM fine-tuning pipeline (QLoRA + GGUF)
Start Llama 70B FP16 · 140 GB ① Quantize AWQ INT4 140 GB → 35 GB ② QLoRA Train adapters ~50 MB added ③ Merge+GGUF Export for Ollama llama.cpp ready Deploy runs on laptop GPU Result: custom 70B LLM on a 48 GB GPU workstation. Fine-tuning cost: 1 GPU × 1–2 days vs 100s of GPUs for full fine-tune.

Case Study B — On-device vision model (Apple Neural Engine)

This is the most aggressive pipeline, used when you need a model to run on a battery-powered chip in under 5 milliseconds with no server connection.

Figure 9.2 — On-device mobile pipeline (NAS + distillation + pruning + quantization)
Start Large cloud vision model ① NAS Find arch for Neural Engine ② Distillation Cloud model teaches student ③ Pruning Structured, iterative ④ INT8 + weight sharing Core ML export Result: runs in <5ms on Neural Engine, on battery, no server. ~80× smaller, ~2–3% accuracy loss.

What each technique targets — the master picture

Figure 9.3 — Each technique attacks a different dimension
Distillation Shrinks the model class ~10× smaller NAS Shapes the architecture chip-optimal Pruning Removes redundancy ~2× smaller Quantization Reduces precision 4–8× smaller Early exit Adapts to input difficulty dynamic saving Combined: 10× × 2× × 4–8× = 80–160× smaller · ~2–5% accuracy loss

The three universal rules

  • Order is not arbitrary. Architecture decisions (NAS, distillation) must come first because they define what you're compressing. Quantization must come last — it is the most destructive step and should be applied to a network already well-optimized by earlier stages.
  • Recover between lossy steps. After any pruning step, fine-tune before the next step. Errors from two lossy operations compound silently if you don't recover in between.
  • Each technique attacks a different dimension. Distillation, NAS, pruning, quantization, early exit, and LoRA address orthogonal problems. None is a substitute for the others — they multiply when stacked correctly.
The gap explained
In 2017, large neural networks ran only on expensive servers. Today they run on your laptop. That entire gap — roughly 100× in efficiency — is almost entirely explained by these techniques stacked together.
References

Foundational Papers & Further Reading

Knowledge Distillation

  1. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. NeurIPS Deep Learning Workshop. — The original paper introducing soft labels, temperature, and dark knowledge. arxiv.org/abs/1503.02531
  2. Romero, A. et al. (2015). FitNets: Hints for Thin Deep Nets. ICLR 2015. — Extends distillation to intermediate layer "hints," not just final outputs. arxiv.org/abs/1412.6550
  3. Sanh, V. et al. (2019). DistilBERT, a distilled version of BERT. NeurIPS EMC² Workshop. — Practical application of distillation to BERT, achieving 60% of its size with 97% of its performance. arxiv.org/abs/1910.01108

Neural Architecture Search (NAS)

  1. Zoph, B. & Le, Q. V. (2017). Neural Architecture Search with Reinforcement Learning. ICLR 2017. — The paper that started the modern NAS era; used RL to search architectures for image classification. arxiv.org/abs/1611.01578
  2. Liu, H., Simonyan, K., & Yang, Y. (2019). DARTS: Differentiable Architecture Search. ICLR 2019. — Made NAS orders of magnitude cheaper by making architecture choices differentiable. arxiv.org/abs/1806.09055
  3. Tan, M. et al. (2019). MnasNet: Platform-Aware Neural Architecture Search for Mobile. CVPR 2019. — Google's evolutionary NAS optimizing for real mobile latency, not FLOPs. arxiv.org/abs/1807.11626
  4. Cai, H. et al. (2020). Once-for-All: Train One Network and Specialize it for Efficient Deployment. ICLR 2020. — Train one supernet, extract free sub-networks for any device. arxiv.org/abs/1908.09791
  5. Cai, H., Zhu, L., & Han, S. (2019). ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. ICLR 2019. — Measures real hardware latency during search rather than using FLOPs as a proxy. arxiv.org/abs/1812.00332

Pruning

  1. Han, S. et al. (2015). Learning both Weights and Connections for Efficient Neural Networks. NeurIPS 2015. — Introduced magnitude-based pruning; showed 9× compression on AlexNet. arxiv.org/abs/1506.02626
  2. Frankle, J. & Carbin, M. (2019). The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. ICLR 2019. — Showed that every large network contains a small "winning" sub-network. arxiv.org/abs/1803.03635
  3. Molchanov, P. et al. (2017). Pruning Convolutional Neural Networks for Resource Efficient Inference. ICLR 2017. — Introduced Taylor expansion-based pruning criterion (gradient × weight). arxiv.org/abs/1611.06440

Quantization

  1. Jacob, B. et al. (2018). Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. CVPR 2018. — Google's foundational paper on quantization-aware training (QAT). arxiv.org/abs/1712.05877
  2. Frantar, E. et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. — Layer-wise quantization with error compensation for LLMs. arxiv.org/abs/2210.17323
  3. Lin, J. et al. (2023). AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. — Identifies and protects the 1% of salient weights via activation statistics. arxiv.org/abs/2306.00978
  4. Gerganov, G. et al. (2023). llama.cpp & GGUF format. GitHub. — Open-source runtime enabling 4-bit quantized LLMs on consumer hardware. github.com/ggerganov/llama.cpp

Low-Rank Factorization & LoRA

  1. Hu, E. et al. (2022). LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. — Introduced trainable low-rank adapter matrices; enabled fine-tuning on a single GPU. arxiv.org/abs/2106.09685
  2. Dettmers, T. et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023. — Combined 4-bit quantization with LoRA; fine-tune a 65B model on a single 48 GB GPU. arxiv.org/abs/2305.14314

Weight Sharing

  1. Han, S., Mao, H., & Dally, W. J. (2016). Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. ICLR 2016 Best Paper. — Combined pruning + weight sharing + Huffman coding for 35× compression. arxiv.org/abs/1510.00149

Early Exit

  1. Teerapittayanon, S., McDanel, B., & Kung, H. T. (2016). BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks. ICPR 2016. — Introduced multi-exit networks with confidence-based early stopping. arxiv.org/abs/1709.01686
  2. Xin, J. et al. (2020). DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference. ACL 2020. — Applied early exit to BERT for NLP, achieving 40% speedup. arxiv.org/abs/2004.12993

Surveys & Broader Reading

  1. Tan, M. & Le, Q. V. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. ICML 2019. — NAS-found architecture that became a widely used baseline for mobile and server models. arxiv.org/abs/1905.11946
  2. Gou, J. et al. (2021). Knowledge Distillation: A Survey. IJCV 2021. — Comprehensive survey covering 40+ distillation variants. arxiv.org/abs/2006.05525
  3. Hoefler, T. et al. (2021). Sparsity in Deep Learning: Pruning and Growth for Efficient Inference and Training in Neural Networks. JMLR 2021. — Definitive survey of pruning methods. arxiv.org/abs/2102.00554