news-impact-v1 / AUDIT_REPORT.md
AylinMaylinn's picture
Add AUDIT_REPORT.md (39,727 bytes)
3ec3d23 verified
|
Raw
History Blame Contribute Delete
39.7 kB
# TBV NewsImpact v1 — Acımasız Denetim Raporu
**Audit date:** 2026-05-13
**Scope:** all claims in the project session brief, against local artifacts and code.
**Coverage:** local pipeline.py (1,731 lines), all v2-v6 runners, FINAL_SESSION_REPORT.md,
news_bbc_guardian.parquet (9.6M rows), 5 rescued chart parquets.
**Out of scope (UNKNOWN):** Drive-hosted .pt checkpoints, attributed_shard_*.parquet,
1,535 chart parquets, pool.json, final_report.json, events log. Torch/transformers
not installed locally, so checkpoint loading and inference smoke tests were
prepared as a script (`audit/audit_inference_skeleton.py`) for the operator to
run on Colab.
---
## 1. Executive Summary
**Verdict: NOT TRUSTED YET.**
The project's narrative — "v6 winner, encoder choice is not the bottleneck, the
data/label noise is the floor" — is *technically true but for the wrong reason*.
There are two simultaneous, mutually-reinforcing failures that fully account for
the observed plateau, and neither is the data:
1. **Cross-attention fusion is mathematically a no-op.** With single-token
query/key/value sequences, the softmax collapses to identity, so the text
encoder output is bit-exactly discarded inside `HybridNewsImpactModel.forward`.
The model is effectively a chart-only classifier wearing a 336M-parameter
text encoder as decoration. (pipeline.py:1119-1124. See proof in
`audit/cross_attention_collapse_proof.md`.)
2. **The "4h" chart dir is half-daily.** `fetch_yfinance_daily` returns DAILY
bars, `fetch_ccxt_4h` returns 4h bars; both are written to `charts_4h/`.
`compute_impact_label`'s 24h pre-window catches at most 1 daily bar →
`pct_change().std()` is NaN → label silently rejected. So labels exist only
for crypto symbols. (pipeline.py:349, 625, 1029-1069.)
These two facts jointly explain every plateau observation in the iteration
scoreboard:
- "v2-v6 cluster at 2.495" → it's the chart-only loss floor.
- "encoder choice doesn't matter" → encoder output is mathematically unused.
- "sig_acc = 0.557 = majority class" → chart-only signal is too weak to separate
significant from non-significant news (because the chart doesn't know whether
the news is bearish, bullish, or unrelated).
- "v6 winner by Δ=0.0004 over v4" → run-to-run noise of effectively the same
chart-only model, not an architecture/encoder effect.
Other notable findings:
3. **36.3% of news rows have null dates.** Cannot be labeled. In the first row
group, every BBC entry (29,467/29,467) had a null date — strong
ingestion-side data corruption.
4. **News source split is BBC 73.8% / Guardian 26.2%**, not the balanced "BBC +
Guardian" framing.
5. **"27 symbol coercions" claim** is off by one — there are 26, and the table
contains a no-op (`^MERV→^MERV`) and a circular pair (`^TPX↔^TOPX`).
6. **`merge_shards_and_test` does naive weight averaging of independently
trained shards without re-evaluating the merged model.** The reported
`mean_val_loss` is the mean of per-shard pre-merge val losses, not the
merged-model performance. The merged model's actual test loss is unknown.
7. **No seed is set anywhere in train_one_shard or runners.** Reproducibility is
non-existent; a Δ of 0.0004 between two runs is not distinguishable from
noise without seed-controlled repeats, which the project did not perform.
8. **No test files exist** for the entire pipeline.
9. **"reason_emb" is a 1024-dim embedding**, not a human-readable explanation.
The product goal "explain why a piece of news matters" requires a decoder,
retrieval scheme, or LLM post-processor — none exists in the current code.
10. **NaN-guard in `train_one_shard` has a subtle scaler bug**: `scaler.update()`
is called after a skipped step without `scaler.scale().backward()` first,
which can corrupt the AMP scaler's internal scale factor.
The architecture and label bugs are root causes; everything else is symptomatic.
---
## 2. Claim Verification Matrix
Categories: A=artifact, B=size/hash, C=arch, D=params, E=frozen, F=ckpt,
M=metric, V=data, L=label, S=split, T=training, I=inference, R=repro,
P=perf-claim, X=prod, O=ops.
| ID | Claim | Cat | Status | Evidence | Confidence | Sev-if-False |
|----|-------|-----|--------|----------|------------|--------------|
| C01 | `pipeline.py` ≈ 1670 lines | A | PARTIAL | 1731 actual (≈4% off) | 1.0 | LOW |
| C02 | pipeline.py syntactically valid | A | PASS | py_compile passed | 1.0 | HIGH |
| C03 | All v2-v6 runner files present locally | A | PASS | All 5 files exist | 1.0 | LOW |
| C04 | `news_bbc_guardian.parquet` ≈ 1.1 GB | B | PASS | 1.07 GiB (1.12 GB) | 1.0 | LOW |
| C05 | News parquet has 9.6M rows | B | PASS | 9,646,340 rows | 1.0 | MED |
| C06 | News source = BBC + Guardian | V | PARTIAL | Yes both, but 73.8/26.2 imbalanced | 1.0 | LOW |
| C07 | `_SYMBOL_COERCIONS` has 27 entries | A | FAIL | 26 entries, includes identity + circular | 1.0 | LOW |
| C08 | 14 delisted symbols skipped | A | PASS | `_DELISTED_SKIP` has 14 entries | 1.0 | LOW |
| C09 | 5 chart files locally rescued | A | PASS | 5 parquets in `rescued/` | 1.0 | LOW |
| C10 | Rescued charts have valid OHLCV | V | PASS | All invariants hold (1 minor low-close violation in URA_F) | 1.0 | MED |
| C11 | All charts are "4h OHLCV" | V | **FAIL** | Yfinance-fetched are daily; only ccxt are 4h. `charts_4h/` is heterogeneous. | 1.0 | **CRITICAL** |
| C12 | `compute_impact_label` produces labels for non-crypto symbols | L | **FAIL** | Daily-bar pre-window yields NaN std → label rejected | 1.0 | **CRITICAL** |
| C13 | Significance defined as |return|/pre_vol > 3.0 | L | PASS | pipeline.py:1068 | 1.0 | n/a |
| C14 | sig threshold is dimensionally consistent | L | FAIL | pre_vol is per-4h-bar std; numerator is 48h return → off by √12 | 0.95 | HIGH |
| C15 | Direction is 3-class up/flat/down | L | PARTIAL | Defined as 3-class but `sign(continuous return)` is almost never 0 → effectively 2-class | 1.0 | MED |
| C16 | HybridNewsImpactModel uses cross-attention to fuse text and chart | C | **FAIL** | Single-token MHA collapses softmax to 1.0 → text contribution = 0 | 1.0 | **CRITICAL** |
| C17 | head_sig dim = 2, head_dir = 3, head_mag = 1, head_reason = 1024 | C | PASS (code) | pipeline.py:1114-1117. `head_reason` output dim = text_dim = 1024 for roberta-large | 1.0 | MED |
| C18 | Fusion is 8-head 1024-dim cross-attention | C | PARTIAL | Configured as 8-head 1024-dim, but functionally degenerate (see C16) | 1.0 | CRITICAL |
| C19 | Chart encoder operates on 128-step 4h OHLCV | C | PARTIAL | chart_window=128 OK; but actual feed is mixed daily/4h (see C11) | 1.0 | HIGH |
| C20 | Trainable params ≈ 6.2M | D | UNKNOWN | Cannot verify w/o torch+ckpt. Skeleton script H2 will measure | n/a | MED |
| C21 | Text encoder is FROZEN in v6 | E | PASS (code) | v6_runner.py:16 + train_one_shard signature wires `freeze_text` | 1.0 | LOW |
| C22 | Only chart_enc + fuse + heads trained in v6 | E | PARTIAL | Code path supports it. But because of C16, text encoder is *also* effectively unused at inference | 1.0 | CRITICAL |
| C23 | final_best.pt exists at Drive path | A | UNKNOWN | No Drive access | n/a | HIGH |
| C24 | final_best.pt is ≈ 1.446 GB | B | UNKNOWN | No Drive access | n/a | LOW |
| C25 | final_best.pt loads strict=True against HybridModel | F | UNKNOWN | Operator must run audit_inference_skeleton.py H1 | n/a | HIGH |
| C26 | Validation loss = 2.4951 reported for v6 | M | PARTIAL | Cannot verify exact value; *but reported value is per-shard val loss averaged across shards, NOT post-merge test loss* (pipeline.py:1482) | 0.95 | HIGH |
| C27 | v6 beats v4 by 0.0004 = meaningful improvement | M | FAIL | Selection logic is `if v6_val < v4_best` w/ no statistical test. No seed set. Within run-to-run noise. | 0.95 | HIGH |
| C28 | Iteration scoreboard plateau ≈ 2.495 ± 0.005 | M | PARTIAL | Cannot verify exact values, but if true it is consistent with the chart-only loss floor — *not* a data/label floor | 0.9 | HIGH |
| C29 | "Data/label noise is the main bottleneck" | M | FAIL | The bottleneck is the architecture (C16) and the daily/4h mismatch (C11/C12). Label noise is also present but is downstream. | 0.95 | HIGH |
| C30 | Train/val split is temporal (last 10%) | S | PASS | pipeline.py:1258-1261 sorts by `t` then slices | 1.0 | n/a |
| C31 | No train/val leakage across shards | S | PARTIAL | Within shard: temporal, OK. Across shards: shard 0 train can overlap in *time* with shard 1's val (interleaved by `iloc[shard::num_shards]`), but per-shard validation is still independent | 0.9 | MED |
| C32 | Cron jobs disabled at session close | O | PASS | FINAL_SESSION_REPORT.md:111 | 1.0 | LOW |
| C33 | Cloudflared tunnel still up | O | UNKNOWN | Local machine state — not directly verifiable | n/a | MED |
| C34 | Vectorized attribution gives 18× speedup (80→4 min) | P | PARTIAL | The current code is vectorized (single big-regex `findall`). Old version not present for direct comparison. 18× ratio not independently verifiable | 0.7 | LOW |
| C35 | Cross-attention seed-controlled training | R | FAIL | No `torch.manual_seed`, `random.seed`, or DataLoader worker seed anywhere in the training pipeline | 1.0 | HIGH |
| C36 | "reason_emb" output is a human-readable explanation | I | FAIL | `head_reason: Linear(text_dim, text_dim)` → 1024-dim float vector. No decoder, no template, no retrieval mechanism in code | 1.0 | HIGH |
| C37 | NaN/inf guard in training is correct | T | PARTIAL | Guard exists (lines 1363-1371) but skips `scaler.scale(loss).backward()` while still calling `scaler.update()` — can corrupt AMP scaler state | 0.95 | MED |
| C38 | Test files exist | X | FAIL | No `test_*.py`, `*_test.py`, or `conftest.py` in project | 1.0 | MED |
| C39 | News dates parse cleanly | V | FAIL | 36.3% of date column is null/unparseable; appears to systematically affect BBC | 1.0 | HIGH |
| C40 | "9 exchanges, 3,159 symbols" pool | V | UNKNOWN | pool.json on Drive, not accessible | n/a | LOW |
| C41 | 1,535 chart files | V | UNKNOWN | charts_4h/ on Drive, not accessible | n/a | LOW |
| C42 | 1,619 crypto fails = DEX-only | V | UNKNOWN | Cannot verify w/o `chart_manifest.json` | n/a | LOW |
| C43 | Inference snippet works as written | I | PARTIAL (untested) | Constructor signature compatible; `state.get('model', state)` should work; but reason_emb is not an explanation as documented | 0.85 | HIGH |
---
## 3. Critical Findings (sorted by severity)
### CRITICAL-1: Cross-Attention Fusion is a Mathematical No-Op
**Where:** `pipeline.py:1119-1124` inside `HybridNewsImpactModel.forward`.
**Evidence:**
```python
fused, _ = self.fuse(t_pool.unsqueeze(1), # (B, 1, D) query
c_pool.unsqueeze(1), # (B, 1, D) key
c_pool.unsqueeze(1)) # (B, 1, D) value
```
With single-token Q/K/V, `softmax` over a 1-element score vector is identically
1.0, so attention weights cannot modulate the value. The output is
`f = W_o · W_v · c_pool + b_o`, a pure linear function of the chart pool. The
text encoder, the W_q projection, the W_k projection — all dead weights at the
forward-pass level. The text encoder's gradients (when not frozen) train only
the encoder's input distribution, never the prediction surface. Full proof in
`audit/cross_attention_collapse_proof.md`.
**Why it matters:** This single bug invalidates every interpretive claim in the
session report. The plateau is the chart-only loss floor; the encoder
indifference is trivially true because the encoder is unused; the v6 "winner"
status is a noise sample of essentially the same chart-only model evaluated
twice. The entire 6.2M-trainable-parameter task layer is being trained to map
a chart pool to (sig, dir, mag, reason) — without news.
**Fix:** Either concatenate-then-project, or have `ChartEncoder.forward` return
a (B, T, D) sequence (don't pool with `.mean(dim=1)`) and use cross-attention
with chart as the *key/value sequence* and text pool as a single query token.
See `audit/cross_attention_collapse_proof.md` § "The fix".
**Recurrence prevention:** Add an automated test:
```python
def test_text_input_affects_output():
out_a = model(tokenize("A"), ..., chart)
out_b = model(tokenize("B"), ..., chart)
assert (out_a['sig_logits'] - out_b['sig_logits']).abs().max() > 1e-6
```
---
### CRITICAL-2: `charts_4h/` Contains Daily Bars; Label Function Silently Rejects Them
**Where:** `pipeline.py:349 (fetch_yfinance_daily)` writes daily bars into
`charts_4h/`; `pipeline.py:1043-1048 (compute_impact_label)` computes
`pre['close'].pct_change().std()` over a 24h pre-window.
**Evidence:**
- The function is named `fetch_yfinance_daily` and uses `yf.download(...)` with
default `interval='1d'`. So one row per market day.
- All 5 locally-rescued files (URA_F, ETHUSD_X, FBHS, XAG_F, idx_IBOV) measured
on disk show `date diff mode = 1 day, median = 1 day, min = 1 day`. Daily, not
4h.
- Only `fetch_ccxt_4h` explicitly requests `'4h'` (line 650).
- `fetch_all_charts` task selector (line 690): non-crypto → `fetch_yfinance_daily`.
- For daily data, a 24h pre-window matches at most 1-2 bars. `pct_change` of a
1-element series returns one NaN; `.dropna().std()` returns NaN. Label is
rejected at line 1051 (`pd.isna(pre_vol)`).
**Why it matters:** The training corpus is *implicitly* filtered to crypto
symbols only. Equities, FX, indices, commodities — all contribute zero labels,
silently. The session's narrative of broad multi-market modeling is not
supported by the data flow.
**Fix:** Either resample yfinance daily to 4h with a documented assumption
(intraday infill is impossible from EOD data — really, use `interval='1h'`
where supported), or shorten `pre_window` to e.g. 3 days when chart resolution
is daily. Make the resolution mismatch loud, not silent: count rejections by
reason in `build_training_pairs`.
**Recurrence prevention:** In `compute_impact_label`, log every rejection
reason; emit a per-shard breakdown of (symbol, n_news, n_labels_produced,
rejection_reasons). Today this information is lost.
---
### CRITICAL-3: Reported `mean_val_loss` is Pre-Merge Per-Shard, Not Merged-Model Performance
**Where:** `pipeline.py:1482` writes `'mean_val_loss': float(np.mean([m['best_val'] for m in shard_models]))` to `final_report.json`. No re-evaluation of the merged model occurs.
**Evidence:** `merge_shards_and_test` takes the mean of per-shard `best_val`
(line 1482), writes the soft-averaged state dict to `final.pt` (line 1474), and
returns. There is no held-out test set used to evaluate the merged model.
**Why it matters:** Soft weight-averaging of independently-trained models often
*degrades* performance — each model has a different loss basin and averaging
weights places the merged point off-basin. The reported "2.4951" describes
*the average performance of four models before they were averaged together*,
not the performance of the model file that gets shipped as `final_best.pt`.
The "winner" comparison between v4 and v6 may therefore not reflect the
actual deployed-model quality.
**Fix:** Reserve a global held-out set up-front (e.g., a temporal hold-out
of the last 5% across all shards, never seen during training). After
`merge_shards_and_test`, evaluate the merged model on this set and report
both numbers: mean-of-shards-val and merged-on-holdout.
---
### HIGH-1: Significance Threshold is Dimensionally Inflated
**Where:** `pipeline.py:1043-1068` in `compute_impact_label`.
**Evidence:**
- `pre_vol = pre['close'].pct_change().std()` over a 24h pre-window of 4h bars
= std of 6 4h-bar returns. Unit: per-4h-bar std.
- `excess = |post_return| / pre_vol` where `post_return = close_post/close_t - 1`
measured over 48h.
- For 48h (=12 bars of 4h), the expected std-equivalent is `pre_vol * √12 ≈
pre_vol * 3.46`. So a 1-σ 48h move has `excess ≈ 3.46`, which already crosses
the threshold of 3.0.
**Why it matters:** The "significant" rate is materially higher than the model
authors intended. Many "ordinary" 48h moves are mis-labeled as significant,
which is consistent with sig_acc plateauing at majority class (~55%): if ~55%
of pairs are labeled significant, predicting "significant always" gets you 55%
accuracy without learning anything.
**Fix:** Use volatility scaling: `excess = |post_return| / (pre_vol * √(post_horizon/bar_size))`.
Or use a separate post-window-matched volatility (rolling 48h std of 4h bars).
Or set threshold dynamically to the 90th percentile of `excess` per symbol.
---
### HIGH-2: News Date Column 36.3% Null; BBC Dates Systematically Missing
**Where:** `news_bbc_guardian.parquet` `date` column (string-typed).
**Evidence:** Full scan over all 10 row groups: 3,504,560 nulls / 9,646,340 =
36.3%. In first row group, BBC had exactly 29,467 entries — identical to the
date-null count in that row group — suggesting every BBC date in rg0 is null.
**Why it matters:** `compute_impact_label` requires a parseable timestamp
(`news_row['date']`, used to select pre/post chart windows). Null-date rows are
unusable. Combined with the daily-bar issue, the effective training corpus is
much smaller than the 9.6M-row corpus claim suggests.
**Fix:** Re-parse the BBC ingestion. If the upstream source provides timestamps,
something is dropping or misparsing them. Until re-parsed, drop null-date rows
loudly at ingestion, not silently at label time.
---
### HIGH-3: No Seed Control, Cannot Distinguish v6 from v4
**Where:** `pipeline.py:train_one_shard` (no `manual_seed` calls). Same for
runners.
**Evidence:** `grep -n "manual_seed\|random.seed\|np.random.seed" pipeline.py
v*_runner.py` returns nothing.
**Why it matters:** Two identical-config runs of a stochastic optimizer with
random data shuffling and dropout will produce different val losses, typically
Δ ≈ 0.005-0.02 for a converged model. The Δ=0.0004 between v4 and v6 is far
inside this noise band. Without seed-controlled triplicate runs, no claim of
"winning" is supportable.
**Fix:** Seed at the top of every runner:
```python
torch.manual_seed(42); random.seed(42); np.random.seed(42)
torch.backends.cudnn.deterministic = True
```
Then run each config 3 times with seeds {42, 1337, 7} and report mean ± std.
---
### HIGH-4: `reason_emb` is Not an Explanation
**Where:** `pipeline.py:1117` `self.head_reason = nn.Linear(text_dim, text_dim)`.
**Evidence:** A 1024-dim float vector is not interpretable. The product spec
("explain why a piece of news matters") needs either:
- A retrieval-based scheme (cosine-match against a labeled bank of rationales).
- A generative decoder (small LM conditioned on `f`).
- A feature-attribution view (e.g., Integrated Gradients of `sig_logits` over
input tokens).
Additionally: the training loss does not appear to supervise `head_reason` at
all. `train_one_shard` only computes loss over `sig_logits`, `dir_logits`, and
`mag`. There is no target embedding being regressed against. So `head_reason`
is trained only by gradient flowing through… no path, because nothing reads it.
The weights of `head_reason` remain at initialization for the entire training.
**Fix:** Either remove the head entirely (it is dead code), or train it. To
train it, supervise it against a meaningful target — e.g., a sentence-encoder
embedding of a per-pair rationale ("oil-sector earnings beat" → encoded with
mpnet), retrieved by manual labeling or a small LLM.
---
### HIGH-5: AMP Scaler State Corruption in NaN-Guard
**Where:** `pipeline.py:1363-1371`.
**Evidence:**
```python
if not torch.isfinite(loss):
nan_skip_count += 1
optim.zero_grad(set_to_none=True)
scaler.update() # <-- called without scaler.scale(loss).backward()
...
continue
```
The recommended pattern in pytorch docs is: if you skip a step, you must NOT
call `scaler.update()` because the scaler's growth tracker counts only valid
steps. Calling `update()` after a skip prematurely grows the loss scale,
making subsequent NaN events more likely.
**Why it matters:** Once you've started accumulating NaN events (which the code
itself notes happen — "50 events at a time"), this guard accelerates the
problem rather than fixes it. Training stability is at risk.
**Fix:** Remove the `scaler.update()` call inside the NaN branch. Pytorch's
recommended pattern:
```python
if not torch.isfinite(loss):
optim.zero_grad(set_to_none=True)
continue # Do NOT touch the scaler.
```
---
### MEDIUM-1: Direction is 3-Class but Effectively 2-Class
**Where:** `pipeline.py:1067` `'direction': int(np.sign(rret))`; then
`pipeline.py:1313` `int(r['direction']) + 1` maps to CE labels {0, 1, 2}.
**Evidence:** `np.sign(continuous_return)` is essentially never exactly 0. So
the "flat" class is never observed. CrossEntropyLoss over a 3-class head with
class 1 (flat) absent biases the head's logits to two-class behavior with a
constant offset for the unused class.
**Why it matters:** Wastes capacity; dir_acc can plateau at the larger of the
up/down class proportions without learning anything.
**Fix:** Either make direction 2-class (up/down) or define "flat" as
`|return| < 0.5 * pre_vol`. Currently dir is conceptually 3-class but
practically 2-class with a deadweight third class.
---
### MEDIUM-2: Symbol Coercion Table is Off-by-One and Sloppy
**Where:** `pipeline.py:307-338`.
**Evidence:**
- Claim: "27-entry _SYMBOL_COERCIONS".
- Actual: 26 entries (counted programmatically).
- One identity: `'^MERV': '^MERV'` (line 311) — does nothing.
- One circular pair: `'^TPX': '^TOPX'` (line 310) and `'^TOPX': '^TPX'` (line 337).
Apply the coercion twice and you cycle.
**Why it matters:** Symptomatic of unreviewed copy-paste edits to a shared
mapping. The circular pair could be a latent bug if any code path applies
coercion iteratively (it doesn't, but the door is open).
**Fix:** Pick one canonical Tokyo Topix symbol and route both ^TPX and ^TOPX to
it. Drop `'^MERV': '^MERV'` (or change to comment that it's an explicit no-op).
Update claim to 27 actual entries.
---
### MEDIUM-3: No Test Files Exist
**Where:** Whole project.
**Evidence:** `find` for `test_*.py`, `*_test.py`, `conftest.py` returns nothing
under `merged_news/`.
**Why it matters:** Future regressions (especially in `compute_impact_label`
and `HybridNewsImpactModel.forward`) will go undetected. The two CRITICAL bugs
above would have been caught by 10 lines of pytest each.
**Fix:** Add `tests/test_pipeline.py` with the recurrence-prevention tests
listed in each CRITICAL finding.
---
### LOW-1: Deprecated APIs
- `torch.cuda.amp.autocast` / `torch.cuda.amp.GradScaler` — deprecated since
pytorch 2.0 in favor of `torch.amp.autocast('cuda', ...)`. Currently emits
DeprecationWarning at runtime.
- `datetime.datetime.utcnow()` (pipeline.py:46, 1484) — deprecated since python
3.12 in favor of `datetime.datetime.now(timezone.utc)`. Currently emits
DeprecationWarning.
Fixes are mechanical and one-line each.
---
## 4. Model Architecture Verification
| Component | Declared | Actual (from code) | Compatible with `final_best.pt`? |
|-----------|----------|--------------------|----------------------------------|
| Text encoder | roberta-large 336M FROZEN | ✓ wired correctly when `freeze_text=True` | UNKNOWN (no ckpt access) |
| Chart encoder | small Transformer over 128-step OHLCV | `ChartEncoder`: in=5, hidden=128, layers=4, nhead=8, out=text_dim. Pools with `mean(dim=1)`. | UNKNOWN |
| Fusion | 8-head cross-attention 1024-dim | `nn.MultiheadAttention(1024, 8, batch_first=True)` — *but collapses to identity-on-value due to T=1* | UNKNOWN |
| `head_sig` | Linear(1024, 2) | ✓ | UNKNOWN |
| `head_dir` | Linear(1024, 3) | ✓ | UNKNOWN |
| `head_mag` | Linear(1024, 1) | ✓ | UNKNOWN |
| `head_reason` | Linear(1024, 1024) | ✓ — **but receives no loss signal**, weights remain at init | UNKNOWN |
| Trainable params | 6.2M | ~6.5M expected (ChartEncoder ≈ 0.6M + fuse ≈ 4.2M + 4 heads ≈ 1.1M). Run script H2 to verify. | UNKNOWN |
The architecture *as written* matches the claims at the layer-construction
level. The collapse bug is a *forward-pass semantics* issue, not a layer-shape
issue. Hence `final_best.pt` likely loads strict=True without errors; the
silent failure is in the data flow, not the weight shapes.
---
## 5. Data & Label Audit
### News distribution
| Property | Value |
|----------|-------|
| Total rows | 9,646,340 |
| Columns | 4 (title, summary, date, source) — all `large_string` |
| Date range | 1999-01-01 → 2026-05-03 |
| Source: The Guardian | 2,525,498 (26.2%) |
| Source: BBC | 7,120,842 (73.8%) |
| Date-null rows | 3,504,560 (36.3%) |
| Date-null rows in row-group 0 | 29,467 = exactly the BBC count in rg0 (suspicious) |
| First-rg duplicate titles | 152,920 / 1,048,576 = 14.6% |
| First-rg duplicate (title, date) | 2,052 (0.2%) |
### Effective training corpus (after losses):
```
9.6M raw news
- 3.5M null-date (-36.3%) → 6.1M
- news that mentions no symbol → unknown
- news mentioning non-crypto symbols → labels rejected by 24h/daily issue
≈ "small fraction of crypto news, before deduplication"
```
### Chart data (5 rescued, only sample available)
| File | Rows | Date range | Resolution | OHLCV invariants |
|------|------|-----------|------------|--------------------|
| ETHUSD_X | 3,107 | 2017-11 → 2026-05 | daily | clean |
| FBHS | 3,107 | 2014-01 → 2026-05 | daily | clean |
| URA_F | 3,107 | 2014-01 → 2026-05 | daily | 1 minor low>close violation |
| XAG_F | 3,107 | 2014-01 → 2026-05 | daily | 219 zero-volume bars |
| idx_IBOV | 3,064 | 2014-01 → 2026-05 | daily | 30 zero-volume bars |
All five are daily, contradicting the directory name `charts_4h/`. Zero-volume
bars for futures (XAG_F) and indices (IBOV) are expected (holidays, low-liquidity
sessions) and not bugs — but `pre_vol` computed across these bars could include
zero-return periods that artificially shrink volatility.
### Label trace: what happens to a typical equity news event?
1. News at 2024-03-15 14:30 UTC: "Apple beats earnings".
2. `attribute_news_to_symbols` matches "Apple" → symbol AAPL.
3. `build_training_pairs` reads `charts_4h/AAPL.parquet` — daily bars.
4. `compute_impact_label` finds:
- pre window (2024-03-14 14:30 → 2024-03-15 14:30): up to 1 daily bar (the
2024-03-15 EOD bar may be tomorrow's close, hence pre window may have 0-1
bars depending on UTC timing of EOD).
- `pre['close'].pct_change().std()` over 0-1 values → NaN.
- `if pd.isna(pre_vol): return {}` → labeled rejected.
5. Pair contributes nothing to training. AAPL effectively contributes nothing.
### Label trace: a crypto news event
1. News at 2024-03-15 14:30 UTC: "Bitcoin surges past $70k".
2. Matches symbol BTC/USDT (via crypto-augmented aliases).
3. Reads ccxt-fetched 4h bars.
4. pre window 24h = 6 bars; pct_change → 5 valid returns; std OK.
5. `excess = |48h return| / 4h-bar-std`.
6. With BTC's typical 4h vol of ~1-2% and 48h returns often in the 2-5% range,
`excess ≈ 1.5-5`. Many will cross threshold 3.0. **Significance rate ~50%
→ majority class**.
This trace matches the observed sig_acc plateau of 0.557 — a chart-only model
on a crypto-heavy corpus with sig threshold inflated by dimensional mismatch.
---
## 6. Metric Audit
### What does the loss decompose to?
`val_loss = ce(sig) + ce(dir) + 0.5 * smoothl1(mag)`
Baseline reference values:
| Baseline | ce(sig) | ce(dir) | 0.5·smoothl1(mag) | total |
|----------|---------|---------|--------------------|-------|
| Random (uniform) | ln 2 ≈ 0.693 | ln 3 ≈ 1.098 | ~0.5 * E[|mag|] (depends on dist) | ≈ 2.0-2.5 |
| Majority class sig (55.7%) | ≈ 0.685 | — | — | — |
| Majority class dir (probably ~50%) | — | ≈ 0.693 | — | — |
| Chart-only model (observed) | ? | ? | ? | 2.495 |
The total of 2.495 is consistent with: majority-class sig (~0.69) + slightly
better than majority dir (~1.0) + a partially trained mag head (~0.8 × 0.5 =
0.4), summed gives ~2.09 — under 2.495. So the model is probably worse than
that decomposition would suggest, or the mag loss dominates. Without
per-component val losses in the report, this can't be pinned down.
### v6 vs v4: is 0.0004 meaningful?
No. Without seed-controlled triplicates:
- Two-sample t-test requires n ≥ 2 per condition with std estimate.
- The session ran each config once. Sample size = 1.
- Run-to-run std of converged validation loss on this kind of model is
typically 0.01-0.05.
- Δ = 0.0004 is two orders of magnitude below the noise floor.
**Verdict: "v6 winner" is a coin flip with extra steps.**
### Suggested baselines that were NOT run
These would distinguish "model learns from news" from "model learns prior":
- **Chart-only ablation** — drop text encoder entirely (or pass random text);
measure val loss. If equal to current v6, the model isn't using text.
(Given C16, this baseline is *guaranteed* equal.)
- **Text-only ablation** — drop chart encoder; measure val loss. If much
higher, news semantics are useful; if equal, news isn't being used.
- **Permuted-label baseline** — shuffle labels within a shard before training.
If val loss drops similarly, the model is learning train/val correlations
unrelated to labels.
- **Permuted-text baseline** — pair each chart with a randomly-assigned news
text from the same shard. If val loss is unchanged from baseline, text is
unused (which we know is the case).
---
## 7. Inference Audit
### What's verifiable from local code
- The inference snippet in the prompt is *almost* correct, but:
- `model.load_state_dict(state['model'] if 'model' in state else state)` — in
`train_one_shard` (line 1437-1443) the checkpoint key is `'state'`, not
`'model'`. In `merge_shards_and_test` (line 1475) it's also `'state'`. So
the snippet should be `state['state'] if 'state' in state else state`.
- `state.get('model', state)` is *forgiving* and won't error — it just falls
through. But the documented-correct line would use 'state'.
### Hypothetical robustness failure modes
Given the cross-attention collapse, all "model robustness" tests degenerate to
"chart robustness". Specifically:
- **Same chart, different news** → output is bit-identical (predicted).
- **Empty/garbled news** → output identical to legitimate news (predicted).
- **Adversarial news ("guaranteed crash")** → output unchanged (predicted).
- **NaN chart input** → likely outputs NaN; no defensive handling in forward.
- **Sub-128-step chart** → in `train_one_shard._DS.__getitem__` lines
1298-1300, the chart is zero-pre-padded if shorter — but at inference time
the caller is responsible for windowing. No safety net.
### Shape contract
The forward returns shapes consistent with the documented contract:
- `sig_logits`: (B, 2) ✓
- `dir_logits`: (B, 3) ✓
- `mag`: (B,) — squeezed from (B,1) ✓
- `reason_emb`: (B, 1024) for roberta-large ✓
The shape contract is OK; the semantic contract (model uses news to predict)
is not.
---
## 8. Code Bug Hunt — Function-Level Summary
| Function | LOC | Status | Notes |
|----------|-----|--------|-------|
| `normalize_date_column` | 5 | OK | Simple wrapper |
| `merge_all_news` | ~40 | OK | Standard concat + dedupe |
| `build_symbol_pool` | ~170 | NOT REVIEWED — UNKNOWN | Large; should be reviewed for pool composition |
| `fetch_yfinance_daily` | ~70 | OK / FAIL-RESOLUTION | Function works; resolution implication is the real bug |
| `fetch_stooq_daily` | ~45 | OK | Fallback path |
| `fetch_crypto_universe_from_exchanges` | ~85 | NOT REVIEWED | |
| `fetch_ccxt_4h` | ~50 | OK | Correctly requests 4h |
| `fetch_all_charts` | ~60 | PARTIAL | Saves daily and 4h into one dir; loud rename suggested |
| `build_symbol_aliases` | ~120 | NOT REVIEWED — UNKNOWN | Long; should be audited for ticker collisions (e.g. "X", "META", "T") |
| `attribute_news_to_symbols` | ~70 | OK | Vectorized findall; correct algorithm. Common-word collision risk depends on `build_symbol_aliases` content. |
| `compute_impact_label` | ~40 | FAIL | C12, C14, C15 above |
| `get_model_classes` | ~60 | FAIL (C16) | Cross-attention collapse |
| `run_data_prep` | ~6 | OK | Orchestrator |
| `run_feature_prep` | ~50 | OK | Shard-then-augment pipeline; `iloc[shard::num_shards]` interleaves correctly |
| `_safe_symbol_filename` | 2 | OK | |
| `build_training_pairs` | ~25 | PARTIAL | Loops one symbol at a time; could be vectorized. Inherits compute_impact_label issues |
| `train_one_shard` | ~225 | PARTIAL | C30 (temporal split) OK, C35 (no seed) FAIL, C37 (scaler bug) FAIL, no head_reason supervision |
| `merge_shards_and_test` | ~40 | FAIL (C26) | Reports pre-merge val, not merged-model test |
| `multi_iterate` | ~110 | PARTIAL | Uses `<` for best-tracking with no statistical confidence (line 1569) |
| `iterate_v2`, `train_remaining_shards_and_merge` | — | NOT REVIEWED | Lower priority for v6 model audit |
---
## 9. Production Readiness
| Dimension | Status | Notes |
|-----------|--------|-------|
| Checkpoint integrity | UNKNOWN | No Drive access |
| Architecture matches spec | FAIL | C16 |
| Reproducibility | FAIL | C35 |
| Tested code | FAIL | C38 |
| Logging / metrics | PARTIAL | Per-step JSONL logging exists, but per-component val loss not logged |
| Dependency pinning | UNKNOWN | No requirements.txt / pyproject in scope |
| Model registry | NONE | Drive-only artifacts |
| Calibration | NONE | No calibration metrics computed |
| Input validation | NONE | No defensive code at inference boundary |
| "Reason" output usability | FAIL | C36 — embedding, not explanation |
| Ops state at session close | PARTIAL | Cron disabled (good), tunnel still up (risk) |
| Security: `torch.load(pt)` | RISK | Pickle deserialization; OK if artifact is trusted |
| Data freshness | PARTIAL | News ends 2026-05-03; close-to-date OK |
**Verdict: not deployable as a "news-impact AI". Deployable, with significant
caveats, as a chart-only crypto move classifier — which is not what's
advertised.**
---
## 10. Recommended Next Actions
### Do immediately (under 1 day)
1. Fix the cross-attention collapse (CRITICAL-1). Either Option A
(concat-then-project) or Option B (chart-as-sequence). Add the
"text-affects-output" pytest. **This is the single highest-ROI fix.**
2. Wire `head_reason` either to a real target or remove it (HIGH-4).
3. Fix the AMP NaN-guard (HIGH-5).
4. Set seeds at the top of every runner; tag the config with seed in the
manifest (HIGH-3).
### Before v7 training
5. Audit and re-derive `compute_impact_label`:
- Pick a single bar resolution (4h or daily) per symbol and *fail loudly* if
mismatched.
- Fix the vol-window ↔ return-window dimensional mismatch (HIGH-1).
- Either drop the "flat" direction class or define it as `|return| < ε·vol`
(MEDIUM-1).
6. Add a held-out global test set; have `merge_shards_and_test` evaluate the
merged model on it (CRITICAL-3).
7. Re-parse BBC news dates (HIGH-2). Confirm ingestion isn't silently dropping
timestamps.
8. Add the 4 baselines from §6: chart-only, text-only, permuted-label,
permuted-text. Report all 5 numbers in the iteration scoreboard.
9. Run each config 3 times with seeds {42, 1337, 7} and report mean ± std.
### Before production deploy
10. Add `tests/test_pipeline.py` with at minimum:
- `test_text_input_affects_output` (CRITICAL-1)
- `test_compute_impact_label_handles_daily_bars` (CRITICAL-2)
- `test_merge_then_evaluate_matches_report` (CRITICAL-3)
- `test_significance_threshold_calibrated` (HIGH-1)
11. Pin dependencies (`requirements.txt` or `uv.lock`).
12. Verify `final_best.pt` loads strict=True; run `audit_inference_skeleton.py`.
13. Decide what `reason_emb` actually does in the product. Either build a
retrieval bank, train a small explanation decoder, or replace with feature
attribution.
14. Add monitoring: log inference latency, input shape, output distribution.
### Nice to have
15. Replace `datetime.utcnow()` and `torch.cuda.amp.*` with their non-deprecated
equivalents (LOW-1).
16. Clean up `_SYMBOL_COERCIONS` (MEDIUM-2).
17. Vectorize `build_training_pairs` (it currently iterates symbol-by-symbol,
row-by-row).
---
## 11. Hypothesis Refutation — Acımasız Mod Sonuçları
The user asked me to test 10 negative hypotheses. Status for each:
| # | Hypothesis | Status | Confidence | Evidence |
|---|-----------|--------|------------|----------|
| 1 | Model only learned class prior, not real signal | **CONFIRMED** | 1.0 | C16 forces this; sig_acc = majority class |
| 2 | Validation split contains leakage | **REFUTED** | 0.9 | Per-shard temporal split is clean; cross-shard temporal overlap is mitigated by per-shard val |
| 3 | Label function is noisy/misaligned | **CONFIRMED** | 1.0 | C12 (daily/4h mismatch), C14 (vol dim mismatch), C15 (flat class) |
| 4 | reason head produces no explanation | **CONFIRMED** | 1.0 | C36 — 1024-dim float vector, no decoder, no training target |
| 5 | v6 winner claim is statistically meaningless | **CONFIRMED** | 1.0 | C27, C35 — no seed, n=1 per config, Δ < noise band |
| 6 | final_best.pt config doesn't match report | UNKNOWN | — | Cannot verify w/o Drive access |
| 7 | Chart input contains future leakage | PARTIAL | 0.7 | Daily-bar timestamps may include a "today's close" that postdates news arrival; window logic is `Date <= t`. For intraday news on daily data, you'd grab same-day EOD which is future. **Suspected leakage for non-crypto.** |
| 8 | News-symbol attribution has high false positives | UNKNOWN | — | Depends on `build_symbol_aliases` content; not audited in this pass |
| 9 | Loss plateau is architecture/label, not capacity | **CONFIRMED** | 0.95 | C16 + C12 + C14 jointly explain the plateau |
| 10 | Production inference snippet is fragile | **CONFIRMED** | 0.9 | Uses wrong checkpoint key (`'model'` instead of `'state'`); falls through by luck. No input validation. |
---
## 12. Answer to "Bu modele güvenebilir miyim?"
1. **`final_best.pt` gerçekten yüklenip inference yapıyor mu?** UNKNOWN
(Drive yok). Mimari shape uyumu var → muhtemelen yükleniyor.
`audit_inference_skeleton.py` ile doğrulayın.
2. **Mimari iddia ile checkpoint uyumlu mu?** Çok büyük olasılıkla evet (shape
düzeyinde). Ama mimari SEMANTİK olarak bozuk (C16).
3. **Raporlanan metrikler doğru ve anlamlı mı?** Hayır. 2.4951, post-merge test
loss değil; pre-merge per-shard val ortalaması. Bu farkı düzeltmeden iddia
"winner" olamaz.
4. **v6 gerçekten daha iyi mi yoksa noise mu?** Noise. Δ=0.0004, seed yok,
sample size 1.
5. **Data/label noise gerçek bottleneck mi?** Kısmen. Label noise gerçek (C12,
C14, C15) ama asıl bottleneck mimari (C16). Mimari düzeltilmeden label'ları
düzeltmek hiçbir şey çözmez.
6. **Major bug/leakage var mı?** Evet, üç kritik bug:
- Cross-attn collapse
- Daily/4h resolution silent rejection
- Pre-merge metric reporting
7. **Production deploy güvenli mi?** Hayır.
8. **v7'den önce mutlaka düzeltilmesi gereken 5 şey:**
1. Cross-attention fusion (CRITICAL-1)
2. Chart resolution mismatch (CRITICAL-2) ve label window
3. Post-merge evaluation (CRITICAL-3)
4. Seed control + triplicate runs
5. `head_reason`'u ya kaldır ya da gerçekten eğit
---
*End of audit report. Generated by Claude on 2026-05-13 from local artifacts.
Drive-bound claims marked UNKNOWN; run `audit/audit_inference_skeleton.py` on
Colab to convert UNKNOWNs to PASS/FAIL.*