--- license: apache-2.0 tags: - object-detection - pytorch - computer-vision - research --- # ObjectModel-v1 ObjectModel-v1 is a clean-room, compact, end-to-end object detector from Bench Labs. It is a research implementation, not a benchmark claim. The model tests whether global semantic reasoning can be compressed into a small fixed latent memory while precise geometry is recovered by query-conditioned sampling from full-resolution pyramid features. The model is NMS-free. It predicts a fixed set of objects and is trained with Hungarian bipartite matching. This is v1: a single from-scratch training run with no pretraining, no hyperparameter search across seeds, and none of the planned ablations run yet. Treat everything below as a first checkpoint in an ongoing process, not a finished result. v2 is expected to do meaningfully better, whether that comes from more training compute, an architecture change informed by v1's ablations, or both. ## Status - Architecture, COCO data path, losses, training, evaluation, profiling, and ONNX export are implemented. - Synthetic forward/loss/backward and data tests are included. - Full COCO2017 training is complete: 100 epochs on a single RTX 5090, peak val AP 0.358 at epoch 95. See [Training Progress](#training-progress) below for the full curve. - No "state of the art" or "beats YOLO/DETR" claim is made. Final AP (0.358) is below the 40-50 range originally set as a competitiveness bar against 20-40M-parameter real-time detectors. This is a real, single-seed, un-pretrained result, not a benchmark claim. See the [Minimum Validation Protocol](#minimum-validation-protocol) for what a competitive claim actually requires: controlled baselines, three seeds, and ablations, none of which have been run yet. ## Training Progress Validation AP by epoch across the full 100-epoch run, rising from near-zero to a peak of 0.358 at epoch 95 | | | |---|---| | Parameters | 40.8M | | Epochs | 100 / 100 (complete) | | Peak val AP (IoU 0.50:0.95) | 0.358 (epoch 95) | | Final val AP (epoch 100) | 0.356 | | Val AP50 / AP75 | 0.542 / 0.377 | | Val AP small / medium / large | 0.184 / 0.385 / 0.493 | | Val AR@100 | 0.573 | | Train loss | 4.11 (from 15.77 at epoch 1) | | Throughput | ~90-110 img/s, batch 32, RTX 5090 | | Single-frame inference | 30.7 ms / 32.5 FPS (batch 1, eager, RTX 5090) | Train loss over the same run, log-shaped as expected for a detector trained from scratch: Train loss by epoch, dropping sharply in the first few epochs then declining gradually from about 6 to 4.1 over the rest of training AP50 (IoU 0.50, a looser localization threshold) against AP75 (IoU 0.75, stricter). Both climb together early on, then AP75 plateaus lower, meaning coarse localization improved faster than precise localization did: AP50 and AP75 by epoch as two lines, AP50 rising to about 0.54 and AP75 to about 0.38 AP broken out by object size. Large objects are detected far more reliably than small ones throughout training, a common pattern for this class of detector and one of the things the required ablations would need to explain: AP by object size (small, medium, large) by epoch as three lines, large objects highest at about 0.49, medium at about 0.39, small lowest at about 0.18 Peak AP came at epoch 95 (0.3578), not the final epoch. That is normal late-training fluctuation, and `best.pt` correctly holds the epoch-95 weights rather than epoch 100's. Against the 40-50 AP range set as the competitiveness bar, this run falls short. It is a real result from a genuine from-scratch run, on the low end of what the project was aiming for. The gap is exactly what the [Minimum Validation Protocol](#minimum-validation-protocol) exists to characterize properly: whether more decoder layers, more latents, longer training, or pretraining would close it is unknown without actually running those ablations. ## Examples Six detections from the final EMA checkpoint (epoch 95, peak AP) on COCO val2017 images, confidence ≥ 0.35. Picked for variety, not cherry-picked for perfection. | | | |---|---| | ![A tray of Japanese desserts and a sake bottle, each item individually boxed](assets/examples/food_tray.jpg) | ![A child on a skateboard with a parent, skateboard detected at high confidence](assets/examples/skatepark.jpg) | | Dessert tray, bottle, bowl, and a dozen individually boxed items | Skate park, parent and child, skateboard at 0.93 confidence | | ![A street market with a person at 0.94 confidence and umbrella awning detected](assets/examples/street_market.jpg) | ![A rodeo scene with multiple people correctly boxed in a crowd](assets/examples/rodeo.jpg) | | Street market, person at 0.94, market umbrella | Rodeo, crowd of people, one animal still mislabeled | | ![A parent holding a baby at a table, with cup, chair, and bottle detected](assets/examples/family.jpg) | ![A rainbow kite in flight above a beach with people and a distant boat](assets/examples/kite_beach.jpg) | | Family at a table, cups, chair, bottle, two people | Beach, kite at 0.80, people along the shore, a distant boat | ## Live Tracking Demo ObjectModel-v1 detections chained through a SORT tracker on real pedestrian footage, boxes and ids tracking people across frames [Full-length video (20s, MP4)](assets/tracking_demo.mp4) ObjectModel-v1 itself has no temporal component. Every frame is detected independently. The clip above chains detections through a from-scratch SORT-style tracker (`src/objectmodel_v1/tracking.py`: constant-velocity Kalman motion model plus IoU/Hungarian frame-to-frame association) to give boxes a persistent id and a short motion trail. Source footage is `vtest.avi`, OpenCV's standard pedestrian test clip (BSD-3, ships with OpenCV), genuine video the model never trained on. This is the same clip run at three points in training, tracker unchanged throughout. Only the detector's checkpoint improved: | Checkpoint | Ids issued over 20s | Longest-lived ids | |---|---|---| | Epoch 13 (AP 0.227) | ~89 | none survive past a few seconds | | Epoch 26 (AP 0.288) | ~84 | 3 ids survive nearly the full clip | | Epoch 95 (AP 0.358, final) | 92 | 4 ids survive nearly the full clip | Track count issued does not fall much, since new people keep entering frame throughout the clip and each one earns a new id, which is correct behavior. Track persistence for people already in frame improved consistently instead. A separate test on a genuinely different scene, an eye-level warehouse clip not shown here, also surfaced a real and distinct limitation: the detector still occasionally hallucinates objects on plain background surfaces, reading a support pillar as "refrigerator", even at this final checkpoint. Tracking quality rides on detection quality, and detection quality on out-of-domain footage (camera angles, lighting, and compression that COCO's photos do not really cover) is visibly weaker than on COCO's own validation images. ## Architecture ```text image -> compact convolutional backbone (strides 8/16/32) -> top-down pyramid fusion -> pooled multi-scale tokens -> fixed latent memory (global semantics) -> learned object queries -> query self-attention -> cross-attention to latent memory -> local sampling around the current query box -> iterative class and box prediction -> object set (no anchors, no NMS) ``` The local sampling radius scales with each query's current width and height. Early decoder layers can search broadly; later layers focus naturally as boxes are refined. During training, an optional dense auxiliary head adds one-to-many spatial supervision. It is discarded for inference and must be evaluated as an ablation, not assumed to help. ## Installation Use Python 3.11 or another PyTorch-supported Python version: ```bash python3.11 -m venv .venv .venv/bin/pip install --upgrade pip .venv/bin/pip install -e '.[coco,export,dev]' ``` For a specific CUDA build, install the matching PyTorch wheel first using the command from , then install ObjectModel-v1. ## Data The default configuration expects COCO 2017: ```text /path/to/coco/ annotations/instances_train2017.json annotations/instances_val2017.json train2017/*.jpg val2017/*.jpg ``` Category IDs are mapped to contiguous training labels and converted back during evaluation. Images without target objects are supported. ## Commands Profile the model before allocating training compute: ```bash objectmodel-profile --config configs/objectmodel_v1.yaml --device cuda ``` Overfit a small dataset first. A full single-GPU command is: ```bash objectmodel-train \ --config configs/objectmodel_v1.yaml \ --data-root /path/to/coco \ --output outputs/objectmodel_v1 ``` Distributed training: ```bash torchrun --standalone --nproc_per_node=8 -m objectmodel_v1.train \ --config configs/objectmodel_v1.yaml \ --data-root /path/to/coco \ --output outputs/objectmodel_v1 ``` Resume and override configuration values: ```bash objectmodel-train \ --config outputs/objectmodel_v1/config.yaml \ --data-root /path/to/coco \ --output outputs/objectmodel_v1 \ --resume outputs/objectmodel_v1/last.pt \ --set train.batch_size=8 ``` Evaluate the EMA checkpoint with canonical `pycocotools` metrics: ```bash objectmodel-eval \ --config outputs/objectmodel_v1/config.yaml \ --checkpoint outputs/objectmodel_v1/best.pt \ --data-root /path/to/coco ``` Export raw logits and normalized `cxcywh` boxes to ONNX: ```bash objectmodel-export \ --config outputs/objectmodel_v1/config.yaml \ --checkpoint outputs/objectmodel_v1/best.pt \ --output outputs/objectmodel_v1/objectmodel-v1.onnx ``` Track detections across video frames (see [Live Tracking Demo](#live-tracking-demo); this is a post-processing layer over independent per-frame detections, not a model capability): ```python from objectmodel_v1.tracking import SortTracker tracker = SortTracker(iou_threshold=0.3, max_age=5, min_hits=2) for frame in video_frames: boxes, labels, scores = detect(frame) # your decode_predictions() call for t in tracker.update(boxes, labels, scores): print(t.id, t.box, t.label, t.score) ``` ## Minimum Validation Protocol Before describing ObjectModel-v1 as competitive, run all models on the same COCO train2017 and val2017 data, image resolution, augmentation budget, training epochs, and hardware. Report: - COCO AP, AP50, AP75, APS, APM, and APL. - Parameters, FLOPs/MACs, FP32/FP16/INT8 artifact sizes. - End-to-end batch-1 median and p95 latency, including preprocessing and decoding. - Peak training and inference memory, GPU-hours, epochs, and images seen. - Three seeds for the principal result, with mean and standard deviation. - Results both from random initialization and with the same permitted pretraining. Required ablations: | Experiment | Question | |---|---| | latent memory vs flattened feature attention | Does compression preserve useful global context? | | local sampler disabled | Does high-resolution geometric evidence improve localization? | | fixed vs box-scaled offsets | Does coarse-to-fine sampling matter? | | dense auxiliary head disabled | Does added supervision improve convergence? | | 1/2/3 latent layers | Where is the accuracy/latency optimum? | | 32/64/96 latents | How aggressively can global context be compressed? | | 3/4/6 decoder layers | What is the anytime speed/accuracy curve? | Suggested external baselines are RT-DETR-R18, D-FINE-N/S, LW-DETR-T/S, and YOLOX-S. Use their official implementations and report their license and measurement setup separately. ## Research Basis ObjectModel-v1 builds on published, independently attributable ideas: - DETR: set prediction and Hungarian matching. - Conditional and Deformable DETR: spatially conditioned/local sparse attention. - RT-DETR: efficient separation of multi-scale encoding and query decoding. - D-FINE: evidence that fine-grained iterative localization is valuable. - DEIM: evidence that one-to-one matching benefits from denser training supervision. - LW-DETR: evidence that compact transformer detectors can compete with real-time CNNs. ObjectModel-v1's specific hypothesis is the combination of a **fixed compressed global memory** and **box-scaled local pyramid sampling**. Publication novelty requires a broader prior-art search and empirical ablations; this repository does not claim that the combination is patent-new. ## License Apache License 2.0. Dataset images, annotations, pretrained weights, and external baselines retain their own licenses and are not included.