How it works#
MouseLite wraps fine-tuned RF-DETR models in a video pipeline: it predicts per frame, links predictions across frames with a multi-object tracker, writes an annotated video, and exports the predictions as a COCO dataset you can re-track or analyse later.
Models#
Three model kinds, fine-tuned on the MTMB dataset:
--kind |
--size |
Output |
|---|---|---|
detection |
nano, small, medium, large |
bounding boxes |
segmentation |
nano, small, medium, large |
boxes + instance masks |
keypoints |
single checkpoint | boxes + pose keypoints |
get_model fetches the matching checkpoint from the Hugging Face Hub on first use and
returns it ready for inference; the download is cached, so it happens once per machine.
Tracking is handled by trackers.
The run path, frame by frame#
Pipeline.run hands a callback to supervision.process_video, which reads the video
with OpenCV and calls the callback once per frame with (frame, frame_idx):
frame ──▶ is frame_idx % every == 0 ?
│ yes │ no
▼ ▼
model.predict(frame, threshold) reuse last detections
│
KeyPoints → Detections (keypoints only)
│
NMS (nms_threshold)
│
top-k by confidence (if top_k set)
│
tracker.update() (skipped when top_k == 1)
│
append COCO image (with frame_index) & annotations
│
└──────────────┬─────────────────────────┘
▼
MetaAnnotator.annotate() → annotated video frame
(+ FPS HUD, + preview window)
Points worth knowing:
everyis a stride, not a skip. Inference and the export only happen on frames whereframe_idx % every == 0, but the annotated video keeps every frame: in-between frames are drawn with the last predictions, and the tracker only ever sees inference frames.- NMS runs after the confidence threshold and before top-k, so
top_k=2means "the two best distinct animals", not two overlapping boxes on one. top_k=1bypasses the tracker entirely. With one animal there is nothing to associate, so every detection is giventracker_id = 0. This is why single-animal runs are indifferent to--tracker.- Keypoints are tracked as boxes — the model's own box, not one fitted around the keypoints — with the keypoints carried alongside so they survive tracking and reach the export.
MetaAnnotator then draws whatever a detection carries: masks, boxes, a #<track_id>
label, keypoints. Colours are looked up by tracker_id, so an animal keeps its colour
across frames. run and retrack share the annotator, so the two videos look alike.
The retrack path#
retrack never touches a model. It reads the source video and the export side by
side, matching frames to COCO images by frame_index, feeds each exported frame's
detections to a fresh tracker, writes <stem>_retracked.mp4, and rewrites track_id
on every annotation in place. Each annotation's score (the model's confidence) is
fed back to the tracker, so confidence-based trackers such as ByteTrack behave as they
would on a live run. Exports written before score was added have none, and every
detection is then seen at 1.0.
everycarries over. If you ran with--every 3, the tracker sees a third of the frames, as it did duringrun; the retracked video still has every frame of the source, at its frame rate, with in-between frames drawn from the last exported one.