File size: 9,752 Bytes
27bbfd1
5b0e13f
 
 
 
27bbfd1
 
5b0e13f
27bbfd1
 
5b0e13f
 
 
 
 
 
 
4d3d37d
 
 
 
5b0e13f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4d3d37d
5b0e13f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
---
title: MILO Benchmark Companion
emoji: πŸ€–
colorFrom: blue
colorTo: purple
sdk: static
pinned: false
license: mit
---

# MILO Benchmark Companion

A read-only companion demo to
[`naishashetty/milo_benchmark`](https://huggingface.co/datasets/naishashetty/milo_benchmark)
on the Hugging Face Hub, and to the
[MILO vision-language-robotics project](https://github.com/NaishaShetty/MILO)
(origin repository). It shows a **leaderboard** and an **episode replay
browser** built entirely from a real benchmark results JSON β€” four
planners (`rule_based`, `behavior_tree`, `htn`, `react` on a local
`qwen2.5:7b` model via Ollama) scored against a fixed 25-task, 5-scene,
3-tier AI2-THOR benchmark.

## This is not a live demo, and that's deliberate

AI2-THOR requires a GPU-backed Unity subprocess to actually simulate a
scene. **HF Spaces' free tier has no GPU and cannot run Unity**, so this
Space cannot execute a real episode interactively, full stop. Rather than
fake it (e.g. replaying a canned animation and implying it's live), this
Space is explicit about being a **static replay** of one specific,
timestamped, already-completed, reproducible benchmark run. If you want
to run the benchmark for real, clone the origin repository and run
`backend/planning_evaluation/run_benchmark.py` yourself against a real
AI2-THOR installation (needs a GPU).

## What's real vs. reconstructed here

This project's own writing convention (see
`experiments/reports/phase_e_milo_benchmark_report.md` in the origin
repo) is to say plainly what's measured vs. illustrative. Applied here:

- **Leaderboard numbers**: real. Read directly from the source results
  JSON's `summary_by_planner` block (never re-derived from `episodes`,
  except for a couple of cost/latency columns not present in
  `summary_by_planner` at all, which are averaged from `episodes`
  instead β€” see `data.py`).
- **Episode replay's logged fields** (planner, scene, instruction,
  goal/object/target, tier, plan_success/execution_success/goal_success,
  wall_clock_ms, failure_cause, llm_retry_attempts): real, read directly
  from that episode's row in the source JSON.
- **Episode replay's "reconstructed plan trace"**: **not** real logged
  data. The source JSON only has aggregate per-episode counts
  (`action_count`, `plan_step_count`) β€” there is no literal per-step
  action log to replay. The step lists shown are built from each task's
  `goal`/`object`/`target` plus this project's documented, deterministic
  planner behavior (`tier3_store` = locate β†’ navigate β†’ pick_up β†’
  navigate β†’ open if openable β†’ place β†’ close if opened, per
  `_deposit()`'s documented behavior in the dataset README and benchmark
  report). Every trace is labeled **"reconstructed plan trace for
  illustration"** in the UI. We chose to build these (rather than the
  simpler, arguably more conservative option of only showing logged
  fields) because a bare goal/object/target/outcome table alone doesn't
  give a reader unfamiliar with the codebase any sense of what a
  planner's "shape" of behavior actually looks like β€” but we did not want
  to under-label them as real trace data, since they aren't.
- **Screenshots**: real UI screenshots from the live MILO product
  (`docs/screenshots/demo/` in the origin repo β€” 3 generic images:
  instruction typed, task in progress, task complete), **not** captured
  per-episode. There is no screenshot for each of the 25 dataset tasks.
  Where shown alongside an episode, the caption says explicitly that it's
  an illustrative example of the UI, not a capture of that literal
  episode's run.

## Why static HTML, not Gradio

This Space started as a Gradio `Blocks` app (a `Dataframe` for the
leaderboard plus a dropdown-driven detail view) β€” a small, well-trodden
pattern for exactly this "table + browsable detail" shape. It was
rebuilt as static HTML/CSS/JS after discovering that **Gradio and Docker
Spaces require an HF Pro subscription for CPU hosting on this account**;
static Spaces are free. Nothing about the page's actual content or logic
needed the change: everything here was already precomputed from a
results JSON with no live Python execution per request, so a static page
loses nothing over the Gradio version for this use case β€” it swaps a
server-rendered `Dataframe`/`Dropdown` for the equivalent plain
HTML table and `<select>`, both populated from the same data at
page-load via `fetch("data.json")`.

## How the static site is built

Nothing here is hand-authored data β€” `data.json` (what `script.js`
fetches) is generated ahead of time by `generate_data.py`, which reuses
`data.py`/`episodes.py`'s exact loading, leaderboard-row, and
plan-trace-reconstruction logic unchanged (only the output target
changed, from an in-process Gradio render to a JSON file). Regenerate it
with:

```bash
python3 generate_data.py
```

whenever a newer `results/milo_benchmark_*.json` lands (see "Data
source" below) β€” then commit and redeploy. `index.html`/`style.css`/
`script.js` never need to change for a data refresh alone.

## Data source and how it stays current

`data.py`'s loader does **not** hardcode a filename β€” it globs
`results/milo_benchmark_*.json` (excluding `*_memory_ablation_*.json`,
a different experiment) and picks the lexicographically-latest
`generated_at_utc`-stamped filename. `results/` currently ships two real
runs from the origin repo's `experiments/results/`:

- `milo_benchmark_20260817T154347Z.json` β€” `react` run against Gemini's
  free tier, which hit a 20-request/day quota after 2 of 25 episodes.
  Kept here for the record, but its `react` numbers are **not** a valid
  capability baseline (see the origin repo's benchmark report, Addendum
  2, for exactly why the raw success rate from that run is misleading).
- `milo_benchmark_20260818T070841Z.json` β€” `react` run against a local
  `qwen2.5:7b` (Ollama), zero external quota dependency, a clean 25/25
  completed run. **This is the file the leaderboard above actually
  loads**, since it's the newer of the two.

If the origin repository produces a newer full run (for example, the
perception-grounded `tier1_locate` check and LLM call/token
instrumentation for `react` landed in the origin repo's
`experiments/results/milo_benchmark_20260818T085629Z.json` after this
Space was first built), copy that JSON into `results/`, re-run
`generate_data.py`, and redeploy. The leaderboard code path is written
defensively for this: if a source JSON lacks a given column (e.g.
`react`'s LLM-call or token counts), that column renders as `β€”` instead
of crashing or silently showing a fabricated `0`.

## Methodology (matches the published dataset card)

`goal_success` is the primary score: it checks **live** post-execution
AI2-THOR object state (`check_goal_live()` in
`backend/planning_evaluation/live_state.py`), not just "did every
planned action dispatch without an error" (`execution_success` β€” a
weaker, separate signal, also shown). Full predicate table and known
limitations (including the honest caveat that `tier1_locate`'s live
check only verifies an object of the named type exists in the scene, not
that it was actually perceived β€” and the separate, additive
`perceived_by_agent` signal that now measures that too) are in
`backend/planning_evaluation/dataset/v1.0/README.md` in the origin
repository β€” read that file for the canonical methodology text; this
Space's copy is a summary, not a re-derivation.

## Links

- Dataset: [huggingface.co/datasets/naishashetty/milo_benchmark](https://huggingface.co/datasets/naishashetty/milo_benchmark)
- Origin repository: [github.com/NaishaShetty/MILO](https://github.com/NaishaShetty/MILO)
- Full benchmark report (methodology, all four planners' real numbers,
  every caveat this README summarizes): `experiments/reports/phase_e_milo_benchmark_report.md`
  in the origin repository.

## Running locally

No server, no dependencies beyond Python (only needed to regenerate
`data.json`, not to view the page):

```bash
python3 -m http.server 8080
```

then open `http://127.0.0.1:8080`. No GPU, no AI2-THOR, no network
access required β€” everything renders from `data.json` and the images in
`screenshots/`.

## Files

```
hf_space/
β”œβ”€β”€ index.html          # page markup: leaderboard / episode replay / about tabs
β”œβ”€β”€ style.css            # styling (light + dark, no build step)
β”œβ”€β”€ script.js             # fetches data.json, renders the leaderboard table and episode detail view
β”œβ”€β”€ data.json              # generated by generate_data.py -- the only thing script.js fetches
β”œβ”€β”€ generate_data.py        # build step: re-derives data.json from results/*.json (not served/run by the Space itself)
β”œβ”€β”€ data.py                  # results JSON loader + leaderboard row builder (defensive re: optional columns) -- used by generate_data.py
β”œβ”€β”€ episodes.py                # curated episode picks + reconstructed-plan-trace builder (clearly labeled) -- used by generate_data.py
β”œβ”€β”€ README.md                   # this file (HF Space card)
β”œβ”€β”€ results/
β”‚   β”œβ”€β”€ milo_benchmark_20260817T154347Z.json   # react/Gemini run (quota-limited, kept for record)
β”‚   └── milo_benchmark_20260818T070841Z.json   # react/qwen2.5:7b run (the one currently loaded)
└── screenshots/
    β”œβ”€β”€ live-01-instruction-typed.png
    β”œβ”€β”€ live-02-task-in-progress.png
    └── live-03-task-complete.png
```

## Status of this Space

Built and verified locally (served via `python3 -m http.server`,
confirmed the page loads, the leaderboard renders real data, and episode
replay works with no console errors) and live on the Hugging Face Hub β€”
see the top of the origin repository's README for the live URL.