---
license: apache-2.0
tags:
- object-detection
- pytorch
- computer-vision
- research
---
# ObjectModel-v1
ObjectModel-v1 is a clean-room, compact, end-to-end object detector from Bench Labs. It is a
research implementation, not a benchmark claim. The model tests
whether global semantic reasoning can be compressed into a small fixed latent memory while
precise geometry is recovered by query-conditioned sampling from full-resolution pyramid
features.
The model is NMS-free. It predicts a fixed set of objects and is trained with Hungarian
bipartite matching.
This is v1: a single from-scratch training run with no pretraining, no hyperparameter search
across seeds, and none of the planned ablations run yet. Treat everything below as a first
checkpoint in an ongoing process, not a finished result. v2 is expected to do meaningfully better,
whether that comes from more training compute, an architecture change informed by v1's ablations,
or both.
## Status
- Architecture, COCO data path, losses, training, evaluation, profiling, and ONNX export are implemented.
- Synthetic forward/loss/backward and data tests are included.
- Full COCO2017 training is complete: 100 epochs on a single RTX 5090, peak val AP 0.358 at epoch 95.
See [Training Progress](#training-progress) below for the full curve.
- No "state of the art" or "beats YOLO/DETR" claim is made. Final AP (0.358) is below the 40-50
range originally set as a competitiveness bar against 20-40M-parameter real-time detectors.
This is a real, single-seed, un-pretrained result, not a benchmark claim. See the
[Minimum Validation Protocol](#minimum-validation-protocol) for what a competitive claim actually
requires: controlled baselines, three seeds, and ablations, none of which have been run yet.
## Training Progress
| | |
|---|---|
| Parameters | 40.8M |
| Epochs | 100 / 100 (complete) |
| Peak val AP (IoU 0.50:0.95) | 0.358 (epoch 95) |
| Final val AP (epoch 100) | 0.356 |
| Val AP50 / AP75 | 0.542 / 0.377 |
| Val AP small / medium / large | 0.184 / 0.385 / 0.493 |
| Val AR@100 | 0.573 |
| Train loss | 4.11 (from 15.77 at epoch 1) |
| Throughput | ~90-110 img/s, batch 32, RTX 5090 |
| Single-frame inference | 30.7 ms / 32.5 FPS (batch 1, eager, RTX 5090) |
Train loss over the same run, log-shaped as expected for a detector trained from scratch:
AP50 (IoU 0.50, a looser localization threshold) against AP75 (IoU 0.75, stricter). Both climb
together early on, then AP75 plateaus lower, meaning coarse localization improved faster than
precise localization did:
AP broken out by object size. Large objects are detected far more reliably than small ones
throughout training, a common pattern for this class of detector and one of the things the
required ablations would need to explain:
Peak AP came at epoch 95 (0.3578), not the final epoch. That is normal late-training fluctuation,
and `best.pt` correctly holds the epoch-95 weights rather than epoch 100's. Against the 40-50 AP
range set as the competitiveness bar, this run falls short. It is a real result from a genuine
from-scratch run, on the low end of what the project was aiming for.
The gap is exactly what the [Minimum Validation Protocol](#minimum-validation-protocol) exists to
characterize properly: whether more decoder layers, more latents, longer training, or pretraining
would close it is unknown without actually running those ablations.
## Examples
Six detections from the final EMA checkpoint (epoch 95, peak AP) on COCO val2017 images,
confidence ≥ 0.35. Picked for variety, not cherry-picked for perfection.
| | |
|---|---|
|  |  |
| Dessert tray, bottle, bowl, and a dozen individually boxed items | Skate park, parent and child, skateboard at 0.93 confidence |
|  |  |
| Street market, person at 0.94, market umbrella | Rodeo, crowd of people, one animal still mislabeled |
|  |  |
| Family at a table, cups, chair, bottle, two people | Beach, kite at 0.80, people along the shore, a distant boat |
## Live Tracking Demo
[Full-length video (20s, MP4)](assets/tracking_demo.mp4)
ObjectModel-v1 itself has no temporal component. Every frame is detected independently. The clip
above chains detections through a from-scratch SORT-style tracker (`src/objectmodel_v1/tracking.py`:
constant-velocity Kalman motion model plus IoU/Hungarian frame-to-frame association) to give boxes
a persistent id and a short motion trail. Source footage is `vtest.avi`, OpenCV's standard
pedestrian test clip (BSD-3, ships with OpenCV), genuine video the model never trained on.
This is the same clip run at three points in training, tracker unchanged throughout. Only the
detector's checkpoint improved:
| Checkpoint | Ids issued over 20s | Longest-lived ids |
|---|---|---|
| Epoch 13 (AP 0.227) | ~89 | none survive past a few seconds |
| Epoch 26 (AP 0.288) | ~84 | 3 ids survive nearly the full clip |
| Epoch 95 (AP 0.358, final) | 92 | 4 ids survive nearly the full clip |
Track count issued does not fall much, since new people keep entering frame throughout the clip
and each one earns a new id, which is correct behavior. Track persistence for people already in
frame improved consistently instead. A separate test on a genuinely different scene, an eye-level
warehouse clip not shown here, also surfaced a real and distinct limitation: the detector still
occasionally hallucinates objects on plain background surfaces, reading a support pillar as
"refrigerator", even at this final checkpoint. Tracking quality rides on detection quality, and
detection quality on out-of-domain footage (camera angles, lighting, and compression that COCO's
photos do not really cover) is visibly weaker than on COCO's own validation images.
## Architecture
```text
image
-> compact convolutional backbone (strides 8/16/32)
-> top-down pyramid fusion
-> pooled multi-scale tokens
-> fixed latent memory (global semantics)
-> learned object queries
-> query self-attention
-> cross-attention to latent memory
-> local sampling around the current query box
-> iterative class and box prediction
-> object set (no anchors, no NMS)
```
The local sampling radius scales with each query's current width and height. Early decoder
layers can search broadly; later layers focus naturally as boxes are refined. During training,
an optional dense auxiliary head adds one-to-many spatial supervision. It is discarded for
inference and must be evaluated as an ablation, not assumed to help.
## Installation
Use Python 3.11 or another PyTorch-supported Python version:
```bash
python3.11 -m venv .venv
.venv/bin/pip install --upgrade pip
.venv/bin/pip install -e '.[coco,export,dev]'
```
For a specific CUDA build, install the matching PyTorch wheel first using the command from
, then install ObjectModel-v1.
## Data
The default configuration expects COCO 2017:
```text
/path/to/coco/
annotations/instances_train2017.json
annotations/instances_val2017.json
train2017/*.jpg
val2017/*.jpg
```
Category IDs are mapped to contiguous training labels and converted back during evaluation.
Images without target objects are supported.
## Commands
Profile the model before allocating training compute:
```bash
objectmodel-profile --config configs/objectmodel_v1.yaml --device cuda
```
Overfit a small dataset first. A full single-GPU command is:
```bash
objectmodel-train \
--config configs/objectmodel_v1.yaml \
--data-root /path/to/coco \
--output outputs/objectmodel_v1
```
Distributed training:
```bash
torchrun --standalone --nproc_per_node=8 -m objectmodel_v1.train \
--config configs/objectmodel_v1.yaml \
--data-root /path/to/coco \
--output outputs/objectmodel_v1
```
Resume and override configuration values:
```bash
objectmodel-train \
--config outputs/objectmodel_v1/config.yaml \
--data-root /path/to/coco \
--output outputs/objectmodel_v1 \
--resume outputs/objectmodel_v1/last.pt \
--set train.batch_size=8
```
Evaluate the EMA checkpoint with canonical `pycocotools` metrics:
```bash
objectmodel-eval \
--config outputs/objectmodel_v1/config.yaml \
--checkpoint outputs/objectmodel_v1/best.pt \
--data-root /path/to/coco
```
Export raw logits and normalized `cxcywh` boxes to ONNX:
```bash
objectmodel-export \
--config outputs/objectmodel_v1/config.yaml \
--checkpoint outputs/objectmodel_v1/best.pt \
--output outputs/objectmodel_v1/objectmodel-v1.onnx
```
Track detections across video frames (see [Live Tracking Demo](#live-tracking-demo); this is a
post-processing layer over independent per-frame detections, not a model capability):
```python
from objectmodel_v1.tracking import SortTracker
tracker = SortTracker(iou_threshold=0.3, max_age=5, min_hits=2)
for frame in video_frames:
boxes, labels, scores = detect(frame) # your decode_predictions() call
for t in tracker.update(boxes, labels, scores):
print(t.id, t.box, t.label, t.score)
```
## Minimum Validation Protocol
Before describing ObjectModel-v1 as competitive, run all models on the same COCO train2017 and
val2017 data, image resolution, augmentation budget, training epochs, and hardware. Report:
- COCO AP, AP50, AP75, APS, APM, and APL.
- Parameters, FLOPs/MACs, FP32/FP16/INT8 artifact sizes.
- End-to-end batch-1 median and p95 latency, including preprocessing and decoding.
- Peak training and inference memory, GPU-hours, epochs, and images seen.
- Three seeds for the principal result, with mean and standard deviation.
- Results both from random initialization and with the same permitted pretraining.
Required ablations:
| Experiment | Question |
|---|---|
| latent memory vs flattened feature attention | Does compression preserve useful global context? |
| local sampler disabled | Does high-resolution geometric evidence improve localization? |
| fixed vs box-scaled offsets | Does coarse-to-fine sampling matter? |
| dense auxiliary head disabled | Does added supervision improve convergence? |
| 1/2/3 latent layers | Where is the accuracy/latency optimum? |
| 32/64/96 latents | How aggressively can global context be compressed? |
| 3/4/6 decoder layers | What is the anytime speed/accuracy curve? |
Suggested external baselines are RT-DETR-R18, D-FINE-N/S, LW-DETR-T/S, and YOLOX-S. Use
their official implementations and report their license and measurement setup separately.
## Research Basis
ObjectModel-v1 builds on published, independently attributable ideas:
- DETR: set prediction and Hungarian matching.
- Conditional and Deformable DETR: spatially conditioned/local sparse attention.
- RT-DETR: efficient separation of multi-scale encoding and query decoding.
- D-FINE: evidence that fine-grained iterative localization is valuable.
- DEIM: evidence that one-to-one matching benefits from denser training supervision.
- LW-DETR: evidence that compact transformer detectors can compete with real-time CNNs.
ObjectModel-v1's specific hypothesis is the combination of a **fixed compressed global memory** and
**box-scaled local pyramid sampling**. Publication novelty requires a broader prior-art search
and empirical ablations; this repository does not claim that the combination is patent-new.
## License
Apache License 2.0. Dataset images, annotations, pretrained weights, and external baselines
retain their own licenses and are not included.