YOLO Model Compression: Distillation, Pruning, Quantization

Knowledge distillation dual-path supervision combining task loss and teacher imitation loss

Key Takeaways

  • An SE attention block adds under 1% extra parameters and leaves the feature-map shape unchanged.
  • Compress in order: distill first, prune second, quantize last, because each step degrades accuracy for the next one.
  • A sample pipeline drops from 40 MB and 11.2 ms to 7 MB and 3.1 ms while mAP50 falls from 0.913 to 0.874.
  • Distillation must freeze the teacher in eval mode under no_grad, or the dark-knowledge source is contaminated.
  • Acceptance testing must be per class, because a 0.8 point overall drop can hide a 6 point loss on a rare class.

YOLO model compression is the last mile between using a detector and rebuilding one: add modules, then shrink the model with distillation, pruning, and quantization. As of 2026, the practical order is distill first, prune second, quantize last, because each step degrades accuracy and later steps should only ever work on the best model available. This guide walks through an SE attention block, a full compression pipeline, and the tracking ecosystem that ships with the same framework.

YOLO practical notes series cover for the model compression and fine-tuning episode

Three Paths to Modifying a Network

By the previous episode you could deliver a detection project end to end. This one covers the last mile from using YOLO to reshaping it: how to change the network structure, how to make a model six times smaller and three times faster, and what else the ecosystem offers beyond detection.

A cold warning first: most “improvements” die from bad experiment design. A paper adds an attention module and gains two points, you reproduce it and gain 0.5, then change the random seed and lose ground. Detection training itself fluctuates by roughly half a point, so a single run cannot separate improvement from luck.

Path Approach Cost and fit
Change config Copy the official YAML, add a P2 small-object head, tune depth and width multipliers No code, reliable gains, first choice
Add modules Write a custom nn.Module and register it in the framework, for example SE Low to medium, the most common paper approach
Swap backbone Plug in an external backbone such as MobileNet or EfficientNet High ceiling but the largest engineering effort, do it last

Four rules for improvement experiments: establish a baseline first, because improvement numbers without a baseline mean nothing; compare at least three random seeds and check whether the means and variance ranges separate; record parameters, GFLOPs, measured latency, and mAP together, since a 0.5 point gain that costs 30% latency is rejected by most businesses; and change one variable at a time.

Hands-On: Adding an SE Attention Block to YOLO

Attention is the most frequently appearing family in improvement papers. The network normally treats every channel of a feature map equally; attention lets important channels speak louder. Squeeze-and-Excitation (SE) is the cleanest implementation — under 1% additional parameters, unchanged shape, and plug-and-play.

Squeeze-and-Excitation attention block showing its squeeze, score, and weight stages

SE has three stages: squeeze, score, and weight. The module pools each channel down to a single number, passes it through a small bottleneck with a SiLU activation and a sigmoid gate, then multiplies the original feature map channel by channel.

import torch
import torch.nn as nn

class SE(nn.Module):
    def __init__(self, ch, reduction=16):
        super().__init__()
        self.pool = nn.AdaptiveAvgPool2d(1)
        self.fc = nn.Sequential(
            nn.Linear(ch, ch // reduction, bias=False),
            nn.SiLU(inplace=True),
            nn.Linear(ch // reduction, ch, bias=False),
            nn.Sigmoid())

    def forward(self, x):
        w = self.pool(x).flatten()              # (N, C), one number per channel
        w = self.fc(w).view(x.shape[0], -1, 1, 1)
        return x * w                            # channel-wise weighting

After writing the module, register it in the framework’s model parser through a zero-intrusion namespace injection, then write a YAML that attaches SE at every Neck scale output, and you can build and train.

from ultralytics.nn import tasks
from se_attention import SE

tasks.parse_model.__globals__["SE"] = SE   # register: YOLO("custom_se.yaml") now works

Pitfall: in testing, the parser does not automatically apply width scaling to custom modules. The channel argument for SE in the YAML must be hand-written as the actual channel count for that scale after scaling — 64 for the n scale rather than 256 — or the build fails or silently computes the wrong shape. Also, after inserting new layers, the source-layer indices of later Concat and Detect blocks must be shifted to match.

The Compression Trio: The Correct Order for Distillation, Pruning, and Quantization

A larger model is more accurate but slower; a smaller one runs but loses accuracy. Compression aims to fit higher accuracy into the same latency budget. The three techniques each handle one segment, and the order matters: distill first, prune second, quantize last. Reverse it and every step is fighting on already-degraded accuracy.

Knowledge distillation lets a small model copy a large model’s homework. Hard labels carry thin information — a “hat” label says little about how hat-like the sample is — while the teacher’s soft distribution (hat 0.82, person 0.11, background 0.07) carries the dark knowledge of inter-class similarity. The student consumes both hard labels and the teacher’s soft output, so the same data yields a denser supervision signal.

Knowledge distillation dual-path supervision combining task loss and teacher imitation loss

Structured pruning slims the network. Train with an L1 sparsity penalty on the batch-normalization scaling factors, let unimportant channels drift toward zero, cut whole channels by global ranking, then fine-tune for a few epochs to recover accuracy. Quantization compresses weights from FP16 to INT8, halving size and lifting throughput by more than 30%. Start with post-training quantization and accept the result if the drop stays within 1 mAP; only consider quantization-aware training if it exceeds budget.

A full pipeline, with example numbers in the range commonly reported by the community:

Stage Model mAP50 Latency Size
Baseline yolo11m FP16 0.913 11.2 ms 40 MB
Distillation yolo11s distilled from m 0.896 5.8 ms 19 MB
Prune 30% + fine-tune pruned s 0.887 4.6 ms 13 MB
INT8 quantization pruned s 0.874 3.1 ms 7 MB

Decision-makers do not look at a single number; they look at which point on the trade-off curve is worth buying.

Pitfall: the two most common compression failures. First, forgetting to freeze the teacher during distillation — if the teacher keeps learning alongside the student, the source of dark knowledge is contaminated, so always run it in eval mode under no_grad. Second, checking only the aggregate score: an overall drop of 0.8 points can hide a 6 point drop on a rare class, which may be exactly the class the business cares about, so acceptance must be checked class by class.

Ecosystem and Tracking: Giving Detections an Identity

The same framework also covers segmentation, pose, classification, and oriented boxes, with a fully symmetric API. Swap the weights to yolo11n-seg.pt for segmentation or a -pose model for pose estimation, and the training, export, and deployment workflow from earlier episodes applies unchanged.

More interesting is multi-object tracking. Detection is stateless and independent per frame, while tracking maintains object identity on top of detection results — you simply switch predict to track.

from ultralytics import YOLO

model = YOLO("yolo11n.pt")
results = model.track(source="video.mp4", stream=True, persist=True, classes=[0])
for r in results:
    ids = r.boxes.id            # may be None on frames where a target appears or disappears
    if ids is not None:
        print("objects:", ids.int().tolist())

The underlying ByteTrack integration has a clever idea: traditional trackers only match high-confidence detections to trajectories, but ByteTrack also uses low-confidence detections. It runs a first high-score matching pass, then a second pass matching leftover trajectories against low-score detections. Low-score detections are often real targets that are half occluded, and recovering them markedly improves continuity in crowded scenes. With stable IDs, line-crossing counts, dwell time in a region, and loitering detection all follow naturally.

Series Wrap-Up: A Capability Self-Check

Capability From Self-check question
Formats and labeling EP.02 / 04 Can you design a labeling spec and build a data pipeline?
Training and tuning EP.02 / 03 Can you read curves, locate the problem, and iterate by priority?
Fundamentals EP.03 Can you hand-write NMS and CIoU and explain DFL and sample assignment?
Deployment and serving EP.04 Can you export correctly and stand up a monitored API?
Improvement and compression EP.05 Can you validate improvements with discipline and compress in the right order?
Ecosystem transfer EP.05 Can you move detection experience to segmentation, pose, or tracking?

If you can answer yes to all six rows, you are no longer just using YOLO — you are doing computer vision engineering. Tools get replaced; these six capabilities do not. When the next generation of models arrives, you will read the documentation and get on board.

Series recap: EP.01 lays the foundation (boxes, IoU, one-stage detectors), EP.02 gets you running (environment, inference, labeling, training), EP.03 opens the hood (architecture, loss, NMS, mAP), EP.04 does real work (the full project flow), and EP.05 moves upward (improvement, compression, multi-task). Read all five and set out with confidence.

Have questions about this article? Feel free to contact us at [email protected] — we’re happy to help!

Frequently Asked Questions

What is the right order for distillation, pruning, and quantization?

Distill first, prune second, and quantize last. Distillation transfers knowledge to a smaller model while accuracy is still high, pruning then removes channels from that model, and quantization compresses the pruned weights to INT8. Reversing the order means every later step works on degraded accuracy.

Why does knowledge distillation need a frozen teacher?

The teacher supplies the soft distribution that carries inter-class dark knowledge. If the teacher keeps training alongside the student, that reference keeps shifting and the supervision signal becomes noisy. Freeze the teacher and run it in eval mode under no_grad.

How much accuracy does INT8 quantization cost?

In the example pipeline, INT8 quantization moved mAP50 from 0.887 to 0.874 while size dropped from 13 MB to 7 MB and latency from 4.6 ms to 3.1 ms. Accept the result if the drop stays within about 1 mAP, but always check per-class numbers.

What is the SE attention block and what does it cost?

Squeeze-and-Excitation pools each channel to one value, scores it through a small bottleneck with a sigmoid gate, then reweights the feature map channel by channel. It adds under 1% extra parameters and leaves the tensor shape unchanged, so it can be inserted almost anywhere.

How does ByteTrack improve crowded-scene tracking?

Classic trackers only match high-confidence detections to existing trajectories. ByteTrack runs a second association pass using low-confidence detections, which often correspond to partially occluded real targets. Recovering them keeps identities stable and noticeably improves continuity in crowded scenes.

About Aomway

Aomway is a technology company specializing in drone and FPV equipment, publishing in-depth analyses of motor control, embedded hardware, and open-source robotics projects. With over 15 years of industry experience, Aomway covers the full spectrum from flight controllers and FPV goggles to thermal imaging cameras and long-range datalinks for the global drone community.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top