Knowledge distillation is a technique where a small model (the student) learns by imitating a large, accurate model (the teacher), rather than learning directly from raw labeled data alone. The goal is to transfer the teacher's "knowledge" into a compact model that is cheaper to run.
Hard labels are human-annotated ground truth: one class is 100%, all others are 0%. They carry no information about similarity between classes.
Soft labels are the probability distributions output by the teacher model (e.g. cat 85%, fox 12%, dog 3%). They capture which classes resemble each other — information invisible in hard labels. Hinton called this dark knowledge.
NAS automates the design of neural network architectures. Instead of a human expert manually choosing layers, layer types, and connections, an algorithm searches for the best combination automatically.
| Method | Search cost | Hardware-aware? | Best for |
|---|---|---|---|
| DARTS | Low (GPU-hours) | Indirect (FLOPs) | Fast prototyping |
| Once-for-All | High once, free after | Yes — many devices | Multi-device deployment |
| ProxylessNAS | Medium | Yes — real latency | Single target chip |
| Evolutionary | Very high (GPU-days) | Yes — real latency | Production models |
Pruning removes parts of a trained neural network that contribute little to its output. 50–90% of weights in a typical trained network can be set to zero with negligible accuracy loss. The Lottery Ticket Hypothesis (2019) showed that inside every large network there is a small "winning sub-network" that does most of the real work.
Quantization reduces numerical precision of weights from 32-bit floats to 8-bit or 4-bit integers. Neural networks are surprisingly tolerant of imprecision — rounding 0.37219184 to an integer like 47 produces nearly identical outputs.
| Format | Full name | Key insight | Best for |
|---|---|---|---|
| GPTQ | Generalized Post-Training Quantization | Compensate rounding error in remaining weights layer by layer | NVIDIA GPU deployment |
| AWQ | Activation-aware Weight Quantization | Protect the 1% of salient weights; aggressively compress the rest | NVIDIA GPU, highest accuracy |
| GGUF | Universal model file for llama.cpp (successor to GGML) | Self-contained file for any hardware — CPU, GPU, Apple Silicon | Laptops, Ollama, llama.cpp |
A weight matrix W (e.g. 1000×1000 = 1,000,000 parameters) can be approximated by two thin matrices A (1000×r) and B (r×1000). At rank 4, A+B store only 8,000 parameters — 125× fewer.
Weight sharing makes multiple weights share the same stored value. Instead of storing each weight as its own 32-bit number, you assign every weight to a cluster and store one representative value per cluster in a codebook. Each weight is replaced by a small index pointing to its cluster.
In production, compression techniques are always combined. Each addresses a different dimension: distillation shrinks architecture, pruning removes redundancy, quantization reduces precision, LoRA enables efficient customization.
Architecture-level decisions first. NAS and distillation define the shape and size of the model. You want to establish this foundation before applying lossy operations like pruning or quantization.
Quantization last. It is the most aggressive single-step transformation. Applying it to a model that has already been well-optimized by pruning and fine-tuning minimizes the total accuracy loss.
Recover between steps. After pruning and before quantizing, always fine-tune to recover lost accuracy. Stacking two lossy steps without recovery compounds the damage significantly.
Early exit adds off-ramps at intermediate layers of a neural network. Instead of every input running through all layers, easy inputs exit after fewer layers while hard inputs continue to the end. The key insight: not all inputs are equally difficult, so why spend the same compute on all of them?
Each exit point has a small classifier head — usually just one or two layers — attached to the main network. After processing through a block of layers, this head outputs a probability distribution. If the top predicted class has probability above a threshold (e.g. 0.90), the network is confident enough to exit now. If the distribution is flat ("20% cat, 18% dog, 15% fox..."), confidence is low and the input continues deeper.
The threshold is a tunable dial: higher = fewer exits (more accurate), lower = more exits (faster). You choose it based on how much accuracy you're willing to trade for speed.
Every other method — pruning, quantization, distillation — produces a fixed speedup regardless of input. Early exit is fundamentally different: its savings are dynamic. The harder the batch, the less you save. This makes it complementary to the other techniques rather than a replacement — you apply pruning and quantization first for a baseline speedup, then layer early exit on top for additional input-adaptive savings.
The biggest mistake practitioners make is picking a technique because it sounds interesting rather than because it matches their actual problem. The table below maps your situation to the right starting point.
| Your situation | Primary technique | Stack with |
|---|---|---|
| Need to run a large model on a laptop / phone | Quantization (INT4) — GGUF / AWQ | Pruning for further size |
| Want to fine-tune a large LLM on your own data | LoRA / QLoRA — adapter-based fine-tuning | INT4 base + distillation |
| Building a model for a specific chip / device | NAS (hardware-aware) — ProxylessNAS / OFA | Distillation + INT8 |
| Large model works well, want a smaller one | Knowledge distillation — teacher → student | Pruning + quantization |
| Model is too slow, accuracy is fine | Structured pruning — remove whole neurons | Quantization after |
| Workload has many "easy" inputs | Early exit — dynamic compute saving | Any other technique first |
| Memory is the bottleneck, not speed | Weight sharing — codebook compression | Low-rank factorization |
This is the most common real-world scenario for practitioners today. You don't train from scratch — that costs millions. Instead:
This is the most aggressive pipeline, used when you need a model to run on a battery-powered chip in under 5 milliseconds with no server connection.