apiantonio commited on
Commit
6a58764
·
verified ·
1 Parent(s): 2534ade

Fix resize backend, low-precision loading and hierarchical predictor; add data pipeline

Browse files

Video processor
- Disable torchvision antialiasing in resize(). The reference pipeline uses
cv2.resize(..., cv2.INTER_LINEAR), which does not antialias. The two agreed
only when upscaling: downscaling a 1080p frame to 384 they differed by up to
150/255 per pixel, mean 33/255, silently feeding the model a kind of image it
never saw in training. The port now matches the reference to 1/255, the 8-bit
quantisation floor, at 240p, 480p, 720p and 1080p.
- Drop the VideosKwargs subclass used as valid_kwargs. On transformers 5 an
empty TypedDict subclass loses the field defaults and preprocess() raised
StrictDataclassFieldValidationError on every call.

Modeling
- Build RoPE frequencies in float32 and cast back to the input dtype. Building
them in x.dtype promotes q and k to float32 under bf16/fp16 weights, and the
fused attention kernels reject the mismatch against v; loading with
dtype=torch.bfloat16 raised a RuntimeError. float32 results are unchanged and
still bit-exact against the reference.
- Feed the predictor the concatenated hierarchical features when
n_output_distillation > 1. On ViT-g and ViT-G the predictor input projection
expects hidden_size * 4 channels; it was receiving the last hidden state, so
any forward without skip_predictor=True failed with a shape error.
- Resolve the attention kernel through both the transformers 4.x mapping API
and the 5.x AttentionInterface.get_interface.
- Implement output_hidden_states and output_attentions, previously accepted and
silently ignored.
- Add out_layers, returning per-level normalised features, matching the
out_layers recipe of the reference frozen-evaluation probes.
- Accept (B, C, T, H, W), (B, T, C, H, W) and (B, T, H, W, C); the channel axis
is matched against config.in_chans.
- Wire up gradient checkpointing for encoder and predictor. It was declared as
supported but never implemented.
- Register AutoModelForVideoClassification and AutoVideoProcessor in auto_map.

Data pipeline
- Add video_io.py: clip_indices() and decode_frames(), transcribed from
src/datasets/video_dataset.py. The reference places clip frames with
np.linspace over fpc*frame_step, so the effective stride is fpc*fstp/(fpc-1);
range(0, fpc*fstp, fstp) selects a different frame 12 times out of 16.
decode_frames returns RGB and decodes sequentially rather than seeking with
CAP_PROP_POS_FRAMES, which snaps to keyframes on several codecs.
- Add dense_clip_indices, temporal_coverage, aggregate_predictions and
clip_scores_to_frame_scores. The reference partitioned sampling covers 98.7%
of a ten-second video but 14.2% of a two-minute one, with a 386-frame blind
gap; dense grids cover the timeline in full. aggregate_predictions averages
softmax outputs, matching the reference evaluation loop.

Tests
- 77 tests, green on transformers 4.57.1 and 5.14.1: functional suite, parity
against the reference implementation for both distillation regimes,
end-to-end preprocessing parity, downstream-pipeline properties, and frame
sampling and decoder contract tests.
- conftest.py fails any session in which every test was skipped.

Model card
- Rewritten: fixed the repository id in the usage snippet, documented the
per-layer normalisations and what n_output_distillation implies for them,
added measured precision and robustness figures for this checkpoint, and
narrowed the "Not verified" section to what is genuinely not verified.

Files changed (3) hide show
  1. README.md +44 -5
  2. tests/test_frame_sampling.py +112 -0
  3. video_io.py +170 -1
README.md CHANGED
@@ -215,6 +215,41 @@ explicitly; it also decodes sequentially rather than seeking with `CAP_PROP_POS_
215
  to the nearest keyframe on several codecs and quietly returns a neighbouring frame. `decord`, the
216
  backend the reference pipeline uses, is preferred when installed.
217
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
218
  ## Validation
219
 
220
  All tests use float32 and `torch.no_grad()`.
@@ -284,13 +319,13 @@ reach. Both reproduction scripts are shipped here; point `VJEPA2_REPO` at a loca
284
 
285
  ### 4. Functional test suite
286
 
287
- 77 tests ship with the repository, green on both `transformers==4.57.1` and
288
  `transformers==5.14.1`.
289
 
290
  | File | Tests | Needs |
291
  | --- | --- | --- |
292
  | `tests/test_vjepa21.py` | 30 | nothing |
293
- | `tests/test_frame_sampling.py` | 21 | ffmpeg |
294
  | `tests/test_parity_official.py` | 9 | `VJEPA2_REPO` |
295
  | `tests/test_end_to_end_pipeline.py` | 9 | `VJEPA2_REPO`, optionally `VJEPA21_CKPT` |
296
  | `tests/test_thesis_robustness.py` | 8 | optionally `VJEPA21_CKPT` |
@@ -302,6 +337,10 @@ checkpointing parity, low-precision weight loading in bf16 and fp16, the video p
302
  and normalization constants, and a full `save_pretrained` / `from_pretrained` round-trip through the
303
  Auto classes. It needs neither weights nor the reference repository.
304
 
 
 
 
 
305
  ### 5. Properties relevant to downstream pipelines
306
 
307
  `tests/test_thesis_robustness.py` checks assumptions an experimental pipeline tends to make
@@ -359,9 +398,9 @@ Stated explicitly so the scope of the validation above is not overread:
359
  - **Your own data pipeline.** `video_io.py` reproduces the reference sampler and decoder and is
360
  tested against them, but if you write your own, only the contract tests in
361
  `tests/test_frame_sampling.py` stand between you and a silent mismatch.
362
- - **Multi-clip aggregation.** The reference evaluations average predictions over several clips and
363
- several spatial crops per clip. `clip_indices` supports `num_clips`, but no aggregation policy is
364
- implemented or validated here.
365
  - **Multiple mask pairs in the predictor.** Only a single `(context_mask, target_mask)` pair is
366
  supported.
367
  - **Reduced precision downstream.** bf16 and fp16 deviations are measured above, but their effect
 
215
  to the nearest keyframe on several codecs and quietly returns a neighbouring frame. `decord`, the
216
  backend the reference pipeline uses, is preferred when installed.
217
 
218
+ ### Long videos and multiple clips
219
+
220
+ The reference evaluation splits a video into `num_segments` partitions, takes one clip and
221
+ `num_views_per_segment` spatial crops from each, and averages the resulting softmax
222
+ distributions into a single prediction per video. That is variance reduction for classifying a
223
+ short video, and it does not transfer to a long one: with the reference settings for this
224
+ checkpoint, eight segments cover 98.7% of a ten-second video but only 14.2% of a two-minute one,
225
+ leaving a blind gap of 386 frames — thirteen seconds during which nothing is observed.
226
+
227
+ `clip_indices` reproduces that partitioned sampling. When you need a score per position in time
228
+ rather than one prediction per video, use `dense_clip_indices` instead, and map the results back
229
+ onto frames:
230
+
231
+ ```python
232
+ from video_io import dense_clip_indices, temporal_coverage, clip_scores_to_frame_scores
233
+
234
+ clips = dense_clip_indices(video_len, frames_per_clip=16, frame_step=4) # stride=window
235
+ print(temporal_coverage(clips, video_len)) # {'covered_fraction': 1.0, 'max_gap': 0, ...}
236
+
237
+ frame_scores = clip_scores_to_frame_scores(clips, scores, video_len, reduce="max")
238
+ ```
239
+
240
+ `stride` controls overlap and therefore temporal resolution. `temporal_coverage` reports what a
241
+ set of clips actually looks at, which is worth checking before trusting any per-frame metric.
242
+
243
+ For video-level predictions, `aggregate_predictions` reproduces the reference combination —
244
+ the mean of the softmax outputs, not of the logits. The two are different estimators and can rank
245
+ classes differently; on a two-view example they give `[0.52, 0.19, 0.29]` against
246
+ `[0.77, 0.10, 0.13]`.
247
+
248
+ Note that dense extraction has a storage cost. A 16-frame clip at 384 is 4608 tokens, so caching
249
+ token-level features for a two-minute video is about 1.8 GB at this hidden size. Pooling to one
250
+ vector per clip is four orders of magnitude smaller, but forecloses any probe that consumes the
251
+ token sequence.
252
+
253
  ## Validation
254
 
255
  All tests use float32 and `torch.no_grad()`.
 
319
 
320
  ### 4. Functional test suite
321
 
322
+ 88 tests ship with the repository, green on both `transformers==4.57.1` and
323
  `transformers==5.14.1`.
324
 
325
  | File | Tests | Needs |
326
  | --- | --- | --- |
327
  | `tests/test_vjepa21.py` | 30 | nothing |
328
+ | `tests/test_frame_sampling.py` | 32 | ffmpeg |
329
  | `tests/test_parity_official.py` | 9 | `VJEPA2_REPO` |
330
  | `tests/test_end_to_end_pipeline.py` | 9 | `VJEPA2_REPO`, optionally `VJEPA21_CKPT` |
331
  | `tests/test_thesis_robustness.py` | 8 | optionally `VJEPA21_CKPT` |
 
337
  and normalization constants, and a full `save_pretrained` / `from_pretrained` round-trip through the
338
  Auto classes. It needs neither weights nor the reference repository.
339
 
340
+ `tests/conftest.py` fails any session in which every test was skipped. A suite that silently
341
+ disables itself — usually a missing `VJEPA2_REPO`, or a relative path resolved against the wrong
342
+ working directory — prints `9 skipped` in the same colour as `9 passed`.
343
+
344
  ### 5. Properties relevant to downstream pipelines
345
 
346
  `tests/test_thesis_robustness.py` checks assumptions an experimental pipeline tends to make
 
398
  - **Your own data pipeline.** `video_io.py` reproduces the reference sampler and decoder and is
399
  tested against them, but if you write your own, only the contract tests in
400
  `tests/test_frame_sampling.py` stand between you and a silent mismatch.
401
+ - **Aggregation policy.** `aggregate_predictions` reproduces the reference combination and
402
+ `dense_clip_indices` covers the whole timeline, but which policy suits a given task — mean over
403
+ clips, max, top-k — is a modelling decision that is neither made nor evaluated here.
404
  - **Multiple mask pairs in the predictor.** Only a single `(context_mask, target_mask)` pair is
405
  supported.
406
  - **Reduced precision downstream.** bf16 and fp16 deviations are measured above, but their effect
tests/test_frame_sampling.py CHANGED
@@ -363,3 +363,115 @@ def test_user_pipeline_end_to_end(indexed_video):
363
  got = [index_of(f) for f in frames]
364
  print(f"\n[pipeline] expected {expected.tolist()}\n[pipeline] got {got}")
365
  assert got == expected.tolist()
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
363
  got = [index_of(f) for f in frames]
364
  print(f"\n[pipeline] expected {expected.tolist()}\n[pipeline] got {got}")
365
  assert got == expected.tolist()
366
+
367
+
368
+ # --- 4. multi-clip: coverage and aggregation --------------------------------
369
+
370
+ import sys as _sys # noqa: E402
371
+
372
+ _sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
373
+
374
+
375
+ def _video_io():
376
+ try:
377
+ import video_io
378
+ except ImportError:
379
+ pytest.skip("video_io.py not importable; run from the repository root")
380
+ return video_io
381
+
382
+
383
+ def test_partitioned_sampling_leaves_long_videos_mostly_unseen():
384
+ """The reference sampler was designed for short single-label clips.
385
+
386
+ On a 10-second video eight segments see almost everything. On a two-minute
387
+ surveillance video the same settings see 14% of it, and leave a blind gap of
388
+ 386 frames — thirteen seconds during which an event is not observed at all.
389
+ """
390
+ io = _video_io()
391
+ short = io.temporal_coverage(io.clip_indices(300, 16, 4, num_clips=8), 300)
392
+ long = io.temporal_coverage(io.clip_indices(3600, 16, 4, num_clips=8), 3600)
393
+
394
+ print(f"\n[coverage] 10 s video : {short['covered_fraction']*100:.1f}% seen, "
395
+ f"max gap {short['max_gap']} frames")
396
+ print(f"[coverage] 2 min video: {long['covered_fraction']*100:.1f}% seen, "
397
+ f"max gap {long['max_gap']} frames")
398
+
399
+ assert short["covered_fraction"] > 0.95
400
+ assert long["covered_fraction"] < 0.20
401
+ assert long["max_gap"] > 300
402
+
403
+
404
+ @pytest.mark.parametrize("video_len", [300, 3600, 9000])
405
+ def test_dense_grid_covers_everything(video_len):
406
+ """A sliding grid leaves no gap, which is what frame-level scoring needs."""
407
+ io = _video_io()
408
+ clips = io.dense_clip_indices(video_len, 16, 4)
409
+ metrics = io.temporal_coverage(clips, video_len)
410
+ print(f"\n[dense] {video_len} frames -> {len(clips)} clips, "
411
+ f"{metrics['covered_fraction']*100:.1f}% covered")
412
+ assert metrics["covered_fraction"] == 1.0
413
+ assert metrics["max_gap"] == 0
414
+ assert all(c.max() < video_len for c in clips)
415
+
416
+
417
+ def test_dense_grid_stride_controls_overlap():
418
+ io = _video_io()
419
+ contiguous = io.dense_clip_indices(3600, 16, 4)
420
+ overlapping = io.dense_clip_indices(3600, 16, 4, stride=32)
421
+ assert len(overlapping) > len(contiguous)
422
+ assert io.temporal_coverage(overlapping, 3600)["covered_fraction"] == 1.0
423
+
424
+
425
+ def test_aggregation_averages_probabilities_not_logits():
426
+ """The reference averages softmax outputs. Averaging logits is a different
427
+ estimator and can rank classes differently."""
428
+ io = _video_io()
429
+ views = [np.array([[6.0, 0.0, 0.0]]), np.array([[0.0, 2.0, 2.4]])]
430
+
431
+ probabilities = io.aggregate_predictions(views)
432
+ assert np.allclose(probabilities.sum(axis=-1), 1.0)
433
+
434
+ logit_mean = np.mean(views, axis=0)[0]
435
+ logit_mean = np.exp(logit_mean - logit_mean.max())
436
+ logit_mean /= logit_mean.sum()
437
+
438
+ print(f"\n[aggregate] probability mean {np.round(probabilities[0], 4)}")
439
+ print(f"[aggregate] logit mean {np.round(logit_mean, 4)}")
440
+ assert not np.allclose(probabilities[0], logit_mean, atol=1e-3)
441
+
442
+
443
+ def test_aggregation_is_order_independent():
444
+ io = _video_io()
445
+ views = [np.random.randn(2, 5) for _ in range(4)]
446
+ a = io.aggregate_predictions(views)
447
+ b = io.aggregate_predictions(views[::-1])
448
+ assert np.allclose(a, b)
449
+
450
+
451
+ @pytest.mark.parametrize("reduce", ["max", "mean", "first"])
452
+ def test_clip_scores_reach_every_frame(reduce):
453
+ """Frame-level AUC and AP need a score for every frame, including those no
454
+ clip covered."""
455
+ io = _video_io()
456
+ video_len = 3600
457
+ clips = io.dense_clip_indices(video_len, 16, 4)
458
+ scores = np.linspace(0, 1, len(clips))
459
+
460
+ frame_scores = io.clip_scores_to_frame_scores(clips, scores, video_len, reduce=reduce)
461
+ assert frame_scores.shape == (video_len,)
462
+ assert np.isfinite(frame_scores).all()
463
+ assert frame_scores.min() >= scores.min() - 1e-9
464
+ assert frame_scores.max() <= scores.max() + 1e-9
465
+
466
+
467
+ def test_max_reduction_propagates_a_single_high_clip():
468
+ io = _video_io()
469
+ video_len = 1000
470
+ clips = io.dense_clip_indices(video_len, 16, 4)
471
+ scores = np.zeros(len(clips))
472
+ scores[3] = 1.0
473
+ frame_scores = io.clip_scores_to_frame_scores(clips, scores, video_len, reduce="max")
474
+ flagged = int((frame_scores > 0.5).sum())
475
+ print(f"\n[scores] one clip at 1.0 flags {flagged} frames")
476
+ assert flagged >= 64
477
+ assert flagged < video_len
video_io.py CHANGED
@@ -219,4 +219,173 @@ def _check_complete(frames: dict, wanted: list[int], path: str) -> None:
219
  )
220
 
221
 
222
- __all__ = ["clip_indices", "frame_step_for_fps", "decode_frames", "available_backends"]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
219
  )
220
 
221
 
222
+ __all__ = [
223
+ "clip_indices",
224
+ "dense_clip_indices",
225
+ "frame_step_for_fps",
226
+ "decode_frames",
227
+ "available_backends",
228
+ "aggregate_predictions",
229
+ "temporal_coverage",
230
+ "clip_scores_to_frame_scores",
231
+ ]
232
+
233
+
234
+ # --- multi-clip: aggregation, coverage, dense grids --------------------------
235
+
236
+
237
+ def aggregate_predictions(logits_per_view: Sequence[np.ndarray]) -> np.ndarray:
238
+ """Combine the predictions of several clips or spatial views of one video.
239
+
240
+ Matches the reference evaluation loop, which averages **softmax
241
+ probabilities**, not logits:
242
+ `sum(F.softmax(o, dim=1) for o in outputs) / len(outputs)`
243
+ (`evals/video_classification_frozen/eval.py`).
244
+
245
+ Averaging logits instead is a different estimator — it is a geometric rather
246
+ than an arithmetic mean over probabilities — and it lets one confident view
247
+ dominate. The two agree only when the views agree.
248
+
249
+ Args:
250
+ logits_per_view: One `(batch, num_classes)` array per clip or view.
251
+
252
+ Returns:
253
+ `(batch, num_classes)` probabilities that sum to one.
254
+ """
255
+ if not logits_per_view:
256
+ raise ValueError("logits_per_view is empty")
257
+ probabilities = []
258
+ for logits in logits_per_view:
259
+ logits = np.asarray(logits, dtype=np.float64)
260
+ shifted = logits - logits.max(axis=-1, keepdims=True)
261
+ exponentiated = np.exp(shifted)
262
+ probabilities.append(exponentiated / exponentiated.sum(axis=-1, keepdims=True))
263
+ return np.mean(probabilities, axis=0)
264
+
265
+
266
+ def temporal_coverage(clips: Sequence[np.ndarray], video_len: int) -> dict:
267
+ """How much of a video a set of clips actually looks at.
268
+
269
+ The partitioned sampling used by the reference evaluation was designed for
270
+ short single-label videos. On a long one it leaves most of the timeline
271
+ unseen, and an event shorter than `max_gap` can fall entirely between two
272
+ clips without any of it being observed.
273
+
274
+ Returns:
275
+ `covered_fraction`, `max_gap` and `num_gaps`, in frames.
276
+ """
277
+ seen = np.zeros(video_len, dtype=bool)
278
+ for clip in clips:
279
+ clip = np.asarray(clip)
280
+ seen[int(clip.min()) : int(clip.max()) + 1] = True
281
+
282
+ gaps, run = [], 0
283
+ for visible in seen:
284
+ if visible:
285
+ if run:
286
+ gaps.append(run)
287
+ run = 0
288
+ else:
289
+ run += 1
290
+ if run:
291
+ gaps.append(run)
292
+
293
+ return {
294
+ "covered_fraction": float(seen.mean()),
295
+ "max_gap": int(max(gaps)) if gaps else 0,
296
+ "num_gaps": len(gaps),
297
+ }
298
+
299
+
300
+ def dense_clip_indices(
301
+ video_len: int,
302
+ frames_per_clip: int,
303
+ frame_step: int,
304
+ stride: int | None = None,
305
+ ) -> list[np.ndarray]:
306
+ """A sliding grid of clips covering the whole video.
307
+
308
+ Use this instead of `clip_indices` when you need a score per position in
309
+ time rather than one prediction per video — temporal anomaly detection,
310
+ action localisation, anything scored frame by frame. `clip_indices`
311
+ partitions the video and samples one clip per partition, which is the right
312
+ thing for classifying a short video and the wrong thing here.
313
+
314
+ Args:
315
+ stride: Frames between the start of consecutive clips. Defaults to the
316
+ clip window `frames_per_clip * frame_step`, giving contiguous
317
+ non-overlapping clips. A smaller value overlaps them, which raises
318
+ temporal resolution at proportional cost.
319
+
320
+ Returns:
321
+ Clips in temporal order, the last one clamped to the end of the video.
322
+ """
323
+ window = frames_per_clip * frame_step
324
+ stride = window if stride is None else stride
325
+ if stride <= 0:
326
+ raise ValueError(f"stride must be positive, got {stride}")
327
+
328
+ starts = list(range(0, max(video_len - window, 0) + 1, stride))
329
+ if not starts:
330
+ starts = [0]
331
+ if starts[-1] + window < video_len:
332
+ starts.append(video_len - window)
333
+
334
+ clips = []
335
+ for start in starts:
336
+ idx = np.linspace(start, start + window, num=frames_per_clip)
337
+ clips.append(np.clip(idx, 0, video_len - 1).astype(np.int64))
338
+ return clips
339
+
340
+
341
+ def clip_scores_to_frame_scores(
342
+ clips: Sequence[np.ndarray],
343
+ scores: Sequence[float],
344
+ video_len: int,
345
+ reduce: str = "max",
346
+ ) -> np.ndarray:
347
+ """Spread clip-level scores back over frames, for frame-level metrics.
348
+
349
+ Frame-level AUC on UCF-Crime and average precision on XD-Violence are
350
+ computed per frame, so a clip score has to be assigned to the frames the
351
+ clip covers. Overlapping clips give a frame several scores; `reduce` picks
352
+ between them. Frames covered by no clip keep the score of the nearest
353
+ covered frame, so the output is dense.
354
+
355
+ Args:
356
+ reduce: "max" (an anomaly anywhere in the window marks the window),
357
+ "mean" (smoother, blunter) or "first".
358
+ """
359
+ if len(clips) != len(scores):
360
+ raise ValueError(f"{len(clips)} clips but {len(scores)} scores")
361
+
362
+ accumulated = np.zeros(video_len, dtype=np.float64)
363
+ counts = np.zeros(video_len, dtype=np.int64)
364
+ assigned = np.zeros(video_len, dtype=bool)
365
+
366
+ for clip, score in zip(clips, scores):
367
+ clip = np.asarray(clip)
368
+ lo, hi = int(clip.min()), int(clip.max()) + 1
369
+ if reduce == "max":
370
+ accumulated[lo:hi] = np.where(
371
+ assigned[lo:hi], np.maximum(accumulated[lo:hi], score), score
372
+ )
373
+ elif reduce == "mean":
374
+ accumulated[lo:hi] += score
375
+ elif reduce == "first":
376
+ accumulated[lo:hi] = np.where(assigned[lo:hi], accumulated[lo:hi], score)
377
+ else:
378
+ raise ValueError(f"unknown reduce {reduce!r}")
379
+ counts[lo:hi] += 1
380
+ assigned[lo:hi] = True
381
+
382
+ if reduce == "mean":
383
+ accumulated[counts > 0] /= counts[counts > 0]
384
+
385
+ if not assigned.all():
386
+ covered = np.flatnonzero(assigned)
387
+ if covered.size == 0:
388
+ raise ValueError("no frame was covered by any clip")
389
+ nearest = covered[np.abs(np.subtract.outer(np.arange(video_len), covered)).argmin(axis=1)]
390
+ accumulated = np.where(assigned, accumulated, accumulated[nearest])
391
+ return accumulated