# TBV NewsImpact v1 — Acımasız Denetim Raporu **Audit date:** 2026-05-13 **Scope:** all claims in the project session brief, against local artifacts and code. **Coverage:** local pipeline.py (1,731 lines), all v2-v6 runners, FINAL_SESSION_REPORT.md, news_bbc_guardian.parquet (9.6M rows), 5 rescued chart parquets. **Out of scope (UNKNOWN):** Drive-hosted .pt checkpoints, attributed_shard_*.parquet, 1,535 chart parquets, pool.json, final_report.json, events log. Torch/transformers not installed locally, so checkpoint loading and inference smoke tests were prepared as a script (`audit/audit_inference_skeleton.py`) for the operator to run on Colab. --- ## 1. Executive Summary **Verdict: NOT TRUSTED YET.** The project's narrative — "v6 winner, encoder choice is not the bottleneck, the data/label noise is the floor" — is *technically true but for the wrong reason*. There are two simultaneous, mutually-reinforcing failures that fully account for the observed plateau, and neither is the data: 1. **Cross-attention fusion is mathematically a no-op.** With single-token query/key/value sequences, the softmax collapses to identity, so the text encoder output is bit-exactly discarded inside `HybridNewsImpactModel.forward`. The model is effectively a chart-only classifier wearing a 336M-parameter text encoder as decoration. (pipeline.py:1119-1124. See proof in `audit/cross_attention_collapse_proof.md`.) 2. **The "4h" chart dir is half-daily.** `fetch_yfinance_daily` returns DAILY bars, `fetch_ccxt_4h` returns 4h bars; both are written to `charts_4h/`. `compute_impact_label`'s 24h pre-window catches at most 1 daily bar → `pct_change().std()` is NaN → label silently rejected. So labels exist only for crypto symbols. (pipeline.py:349, 625, 1029-1069.) These two facts jointly explain every plateau observation in the iteration scoreboard: - "v2-v6 cluster at 2.495" → it's the chart-only loss floor. - "encoder choice doesn't matter" → encoder output is mathematically unused. - "sig_acc = 0.557 = majority class" → chart-only signal is too weak to separate significant from non-significant news (because the chart doesn't know whether the news is bearish, bullish, or unrelated). - "v6 winner by Δ=0.0004 over v4" → run-to-run noise of effectively the same chart-only model, not an architecture/encoder effect. Other notable findings: 3. **36.3% of news rows have null dates.** Cannot be labeled. In the first row group, every BBC entry (29,467/29,467) had a null date — strong ingestion-side data corruption. 4. **News source split is BBC 73.8% / Guardian 26.2%**, not the balanced "BBC + Guardian" framing. 5. **"27 symbol coercions" claim** is off by one — there are 26, and the table contains a no-op (`^MERV→^MERV`) and a circular pair (`^TPX↔^TOPX`). 6. **`merge_shards_and_test` does naive weight averaging of independently trained shards without re-evaluating the merged model.** The reported `mean_val_loss` is the mean of per-shard pre-merge val losses, not the merged-model performance. The merged model's actual test loss is unknown. 7. **No seed is set anywhere in train_one_shard or runners.** Reproducibility is non-existent; a Δ of 0.0004 between two runs is not distinguishable from noise without seed-controlled repeats, which the project did not perform. 8. **No test files exist** for the entire pipeline. 9. **"reason_emb" is a 1024-dim embedding**, not a human-readable explanation. The product goal "explain why a piece of news matters" requires a decoder, retrieval scheme, or LLM post-processor — none exists in the current code. 10. **NaN-guard in `train_one_shard` has a subtle scaler bug**: `scaler.update()` is called after a skipped step without `scaler.scale().backward()` first, which can corrupt the AMP scaler's internal scale factor. The architecture and label bugs are root causes; everything else is symptomatic. --- ## 2. Claim Verification Matrix Categories: A=artifact, B=size/hash, C=arch, D=params, E=frozen, F=ckpt, M=metric, V=data, L=label, S=split, T=training, I=inference, R=repro, P=perf-claim, X=prod, O=ops. | ID | Claim | Cat | Status | Evidence | Confidence | Sev-if-False | |----|-------|-----|--------|----------|------------|--------------| | C01 | `pipeline.py` ≈ 1670 lines | A | PARTIAL | 1731 actual (≈4% off) | 1.0 | LOW | | C02 | pipeline.py syntactically valid | A | PASS | py_compile passed | 1.0 | HIGH | | C03 | All v2-v6 runner files present locally | A | PASS | All 5 files exist | 1.0 | LOW | | C04 | `news_bbc_guardian.parquet` ≈ 1.1 GB | B | PASS | 1.07 GiB (1.12 GB) | 1.0 | LOW | | C05 | News parquet has 9.6M rows | B | PASS | 9,646,340 rows | 1.0 | MED | | C06 | News source = BBC + Guardian | V | PARTIAL | Yes both, but 73.8/26.2 imbalanced | 1.0 | LOW | | C07 | `_SYMBOL_COERCIONS` has 27 entries | A | FAIL | 26 entries, includes identity + circular | 1.0 | LOW | | C08 | 14 delisted symbols skipped | A | PASS | `_DELISTED_SKIP` has 14 entries | 1.0 | LOW | | C09 | 5 chart files locally rescued | A | PASS | 5 parquets in `rescued/` | 1.0 | LOW | | C10 | Rescued charts have valid OHLCV | V | PASS | All invariants hold (1 minor low-close violation in URA_F) | 1.0 | MED | | C11 | All charts are "4h OHLCV" | V | **FAIL** | Yfinance-fetched are daily; only ccxt are 4h. `charts_4h/` is heterogeneous. | 1.0 | **CRITICAL** | | C12 | `compute_impact_label` produces labels for non-crypto symbols | L | **FAIL** | Daily-bar pre-window yields NaN std → label rejected | 1.0 | **CRITICAL** | | C13 | Significance defined as |return|/pre_vol > 3.0 | L | PASS | pipeline.py:1068 | 1.0 | n/a | | C14 | sig threshold is dimensionally consistent | L | FAIL | pre_vol is per-4h-bar std; numerator is 48h return → off by √12 | 0.95 | HIGH | | C15 | Direction is 3-class up/flat/down | L | PARTIAL | Defined as 3-class but `sign(continuous return)` is almost never 0 → effectively 2-class | 1.0 | MED | | C16 | HybridNewsImpactModel uses cross-attention to fuse text and chart | C | **FAIL** | Single-token MHA collapses softmax to 1.0 → text contribution = 0 | 1.0 | **CRITICAL** | | C17 | head_sig dim = 2, head_dir = 3, head_mag = 1, head_reason = 1024 | C | PASS (code) | pipeline.py:1114-1117. `head_reason` output dim = text_dim = 1024 for roberta-large | 1.0 | MED | | C18 | Fusion is 8-head 1024-dim cross-attention | C | PARTIAL | Configured as 8-head 1024-dim, but functionally degenerate (see C16) | 1.0 | CRITICAL | | C19 | Chart encoder operates on 128-step 4h OHLCV | C | PARTIAL | chart_window=128 OK; but actual feed is mixed daily/4h (see C11) | 1.0 | HIGH | | C20 | Trainable params ≈ 6.2M | D | UNKNOWN | Cannot verify w/o torch+ckpt. Skeleton script H2 will measure | n/a | MED | | C21 | Text encoder is FROZEN in v6 | E | PASS (code) | v6_runner.py:16 + train_one_shard signature wires `freeze_text` | 1.0 | LOW | | C22 | Only chart_enc + fuse + heads trained in v6 | E | PARTIAL | Code path supports it. But because of C16, text encoder is *also* effectively unused at inference | 1.0 | CRITICAL | | C23 | final_best.pt exists at Drive path | A | UNKNOWN | No Drive access | n/a | HIGH | | C24 | final_best.pt is ≈ 1.446 GB | B | UNKNOWN | No Drive access | n/a | LOW | | C25 | final_best.pt loads strict=True against HybridModel | F | UNKNOWN | Operator must run audit_inference_skeleton.py H1 | n/a | HIGH | | C26 | Validation loss = 2.4951 reported for v6 | M | PARTIAL | Cannot verify exact value; *but reported value is per-shard val loss averaged across shards, NOT post-merge test loss* (pipeline.py:1482) | 0.95 | HIGH | | C27 | v6 beats v4 by 0.0004 = meaningful improvement | M | FAIL | Selection logic is `if v6_val < v4_best` w/ no statistical test. No seed set. Within run-to-run noise. | 0.95 | HIGH | | C28 | Iteration scoreboard plateau ≈ 2.495 ± 0.005 | M | PARTIAL | Cannot verify exact values, but if true it is consistent with the chart-only loss floor — *not* a data/label floor | 0.9 | HIGH | | C29 | "Data/label noise is the main bottleneck" | M | FAIL | The bottleneck is the architecture (C16) and the daily/4h mismatch (C11/C12). Label noise is also present but is downstream. | 0.95 | HIGH | | C30 | Train/val split is temporal (last 10%) | S | PASS | pipeline.py:1258-1261 sorts by `t` then slices | 1.0 | n/a | | C31 | No train/val leakage across shards | S | PARTIAL | Within shard: temporal, OK. Across shards: shard 0 train can overlap in *time* with shard 1's val (interleaved by `iloc[shard::num_shards]`), but per-shard validation is still independent | 0.9 | MED | | C32 | Cron jobs disabled at session close | O | PASS | FINAL_SESSION_REPORT.md:111 | 1.0 | LOW | | C33 | Cloudflared tunnel still up | O | UNKNOWN | Local machine state — not directly verifiable | n/a | MED | | C34 | Vectorized attribution gives 18× speedup (80→4 min) | P | PARTIAL | The current code is vectorized (single big-regex `findall`). Old version not present for direct comparison. 18× ratio not independently verifiable | 0.7 | LOW | | C35 | Cross-attention seed-controlled training | R | FAIL | No `torch.manual_seed`, `random.seed`, or DataLoader worker seed anywhere in the training pipeline | 1.0 | HIGH | | C36 | "reason_emb" output is a human-readable explanation | I | FAIL | `head_reason: Linear(text_dim, text_dim)` → 1024-dim float vector. No decoder, no template, no retrieval mechanism in code | 1.0 | HIGH | | C37 | NaN/inf guard in training is correct | T | PARTIAL | Guard exists (lines 1363-1371) but skips `scaler.scale(loss).backward()` while still calling `scaler.update()` — can corrupt AMP scaler state | 0.95 | MED | | C38 | Test files exist | X | FAIL | No `test_*.py`, `*_test.py`, or `conftest.py` in project | 1.0 | MED | | C39 | News dates parse cleanly | V | FAIL | 36.3% of date column is null/unparseable; appears to systematically affect BBC | 1.0 | HIGH | | C40 | "9 exchanges, 3,159 symbols" pool | V | UNKNOWN | pool.json on Drive, not accessible | n/a | LOW | | C41 | 1,535 chart files | V | UNKNOWN | charts_4h/ on Drive, not accessible | n/a | LOW | | C42 | 1,619 crypto fails = DEX-only | V | UNKNOWN | Cannot verify w/o `chart_manifest.json` | n/a | LOW | | C43 | Inference snippet works as written | I | PARTIAL (untested) | Constructor signature compatible; `state.get('model', state)` should work; but reason_emb is not an explanation as documented | 0.85 | HIGH | --- ## 3. Critical Findings (sorted by severity) ### CRITICAL-1: Cross-Attention Fusion is a Mathematical No-Op **Where:** `pipeline.py:1119-1124` inside `HybridNewsImpactModel.forward`. **Evidence:** ```python fused, _ = self.fuse(t_pool.unsqueeze(1), # (B, 1, D) query c_pool.unsqueeze(1), # (B, 1, D) key c_pool.unsqueeze(1)) # (B, 1, D) value ``` With single-token Q/K/V, `softmax` over a 1-element score vector is identically 1.0, so attention weights cannot modulate the value. The output is `f = W_o · W_v · c_pool + b_o`, a pure linear function of the chart pool. The text encoder, the W_q projection, the W_k projection — all dead weights at the forward-pass level. The text encoder's gradients (when not frozen) train only the encoder's input distribution, never the prediction surface. Full proof in `audit/cross_attention_collapse_proof.md`. **Why it matters:** This single bug invalidates every interpretive claim in the session report. The plateau is the chart-only loss floor; the encoder indifference is trivially true because the encoder is unused; the v6 "winner" status is a noise sample of essentially the same chart-only model evaluated twice. The entire 6.2M-trainable-parameter task layer is being trained to map a chart pool to (sig, dir, mag, reason) — without news. **Fix:** Either concatenate-then-project, or have `ChartEncoder.forward` return a (B, T, D) sequence (don't pool with `.mean(dim=1)`) and use cross-attention with chart as the *key/value sequence* and text pool as a single query token. See `audit/cross_attention_collapse_proof.md` § "The fix". **Recurrence prevention:** Add an automated test: ```python def test_text_input_affects_output(): out_a = model(tokenize("A"), ..., chart) out_b = model(tokenize("B"), ..., chart) assert (out_a['sig_logits'] - out_b['sig_logits']).abs().max() > 1e-6 ``` --- ### CRITICAL-2: `charts_4h/` Contains Daily Bars; Label Function Silently Rejects Them **Where:** `pipeline.py:349 (fetch_yfinance_daily)` writes daily bars into `charts_4h/`; `pipeline.py:1043-1048 (compute_impact_label)` computes `pre['close'].pct_change().std()` over a 24h pre-window. **Evidence:** - The function is named `fetch_yfinance_daily` and uses `yf.download(...)` with default `interval='1d'`. So one row per market day. - All 5 locally-rescued files (URA_F, ETHUSD_X, FBHS, XAG_F, idx_IBOV) measured on disk show `date diff mode = 1 day, median = 1 day, min = 1 day`. Daily, not 4h. - Only `fetch_ccxt_4h` explicitly requests `'4h'` (line 650). - `fetch_all_charts` task selector (line 690): non-crypto → `fetch_yfinance_daily`. - For daily data, a 24h pre-window matches at most 1-2 bars. `pct_change` of a 1-element series returns one NaN; `.dropna().std()` returns NaN. Label is rejected at line 1051 (`pd.isna(pre_vol)`). **Why it matters:** The training corpus is *implicitly* filtered to crypto symbols only. Equities, FX, indices, commodities — all contribute zero labels, silently. The session's narrative of broad multi-market modeling is not supported by the data flow. **Fix:** Either resample yfinance daily to 4h with a documented assumption (intraday infill is impossible from EOD data — really, use `interval='1h'` where supported), or shorten `pre_window` to e.g. 3 days when chart resolution is daily. Make the resolution mismatch loud, not silent: count rejections by reason in `build_training_pairs`. **Recurrence prevention:** In `compute_impact_label`, log every rejection reason; emit a per-shard breakdown of (symbol, n_news, n_labels_produced, rejection_reasons). Today this information is lost. --- ### CRITICAL-3: Reported `mean_val_loss` is Pre-Merge Per-Shard, Not Merged-Model Performance **Where:** `pipeline.py:1482` writes `'mean_val_loss': float(np.mean([m['best_val'] for m in shard_models]))` to `final_report.json`. No re-evaluation of the merged model occurs. **Evidence:** `merge_shards_and_test` takes the mean of per-shard `best_val` (line 1482), writes the soft-averaged state dict to `final.pt` (line 1474), and returns. There is no held-out test set used to evaluate the merged model. **Why it matters:** Soft weight-averaging of independently-trained models often *degrades* performance — each model has a different loss basin and averaging weights places the merged point off-basin. The reported "2.4951" describes *the average performance of four models before they were averaged together*, not the performance of the model file that gets shipped as `final_best.pt`. The "winner" comparison between v4 and v6 may therefore not reflect the actual deployed-model quality. **Fix:** Reserve a global held-out set up-front (e.g., a temporal hold-out of the last 5% across all shards, never seen during training). After `merge_shards_and_test`, evaluate the merged model on this set and report both numbers: mean-of-shards-val and merged-on-holdout. --- ### HIGH-1: Significance Threshold is Dimensionally Inflated **Where:** `pipeline.py:1043-1068` in `compute_impact_label`. **Evidence:** - `pre_vol = pre['close'].pct_change().std()` over a 24h pre-window of 4h bars = std of 6 4h-bar returns. Unit: per-4h-bar std. - `excess = |post_return| / pre_vol` where `post_return = close_post/close_t - 1` measured over 48h. - For 48h (=12 bars of 4h), the expected std-equivalent is `pre_vol * √12 ≈ pre_vol * 3.46`. So a 1-σ 48h move has `excess ≈ 3.46`, which already crosses the threshold of 3.0. **Why it matters:** The "significant" rate is materially higher than the model authors intended. Many "ordinary" 48h moves are mis-labeled as significant, which is consistent with sig_acc plateauing at majority class (~55%): if ~55% of pairs are labeled significant, predicting "significant always" gets you 55% accuracy without learning anything. **Fix:** Use volatility scaling: `excess = |post_return| / (pre_vol * √(post_horizon/bar_size))`. Or use a separate post-window-matched volatility (rolling 48h std of 4h bars). Or set threshold dynamically to the 90th percentile of `excess` per symbol. --- ### HIGH-2: News Date Column 36.3% Null; BBC Dates Systematically Missing **Where:** `news_bbc_guardian.parquet` `date` column (string-typed). **Evidence:** Full scan over all 10 row groups: 3,504,560 nulls / 9,646,340 = 36.3%. In first row group, BBC had exactly 29,467 entries — identical to the date-null count in that row group — suggesting every BBC date in rg0 is null. **Why it matters:** `compute_impact_label` requires a parseable timestamp (`news_row['date']`, used to select pre/post chart windows). Null-date rows are unusable. Combined with the daily-bar issue, the effective training corpus is much smaller than the 9.6M-row corpus claim suggests. **Fix:** Re-parse the BBC ingestion. If the upstream source provides timestamps, something is dropping or misparsing them. Until re-parsed, drop null-date rows loudly at ingestion, not silently at label time. --- ### HIGH-3: No Seed Control, Cannot Distinguish v6 from v4 **Where:** `pipeline.py:train_one_shard` (no `manual_seed` calls). Same for runners. **Evidence:** `grep -n "manual_seed\|random.seed\|np.random.seed" pipeline.py v*_runner.py` returns nothing. **Why it matters:** Two identical-config runs of a stochastic optimizer with random data shuffling and dropout will produce different val losses, typically Δ ≈ 0.005-0.02 for a converged model. The Δ=0.0004 between v4 and v6 is far inside this noise band. Without seed-controlled triplicate runs, no claim of "winning" is supportable. **Fix:** Seed at the top of every runner: ```python torch.manual_seed(42); random.seed(42); np.random.seed(42) torch.backends.cudnn.deterministic = True ``` Then run each config 3 times with seeds {42, 1337, 7} and report mean ± std. --- ### HIGH-4: `reason_emb` is Not an Explanation **Where:** `pipeline.py:1117` `self.head_reason = nn.Linear(text_dim, text_dim)`. **Evidence:** A 1024-dim float vector is not interpretable. The product spec ("explain why a piece of news matters") needs either: - A retrieval-based scheme (cosine-match against a labeled bank of rationales). - A generative decoder (small LM conditioned on `f`). - A feature-attribution view (e.g., Integrated Gradients of `sig_logits` over input tokens). Additionally: the training loss does not appear to supervise `head_reason` at all. `train_one_shard` only computes loss over `sig_logits`, `dir_logits`, and `mag`. There is no target embedding being regressed against. So `head_reason` is trained only by gradient flowing through… no path, because nothing reads it. The weights of `head_reason` remain at initialization for the entire training. **Fix:** Either remove the head entirely (it is dead code), or train it. To train it, supervise it against a meaningful target — e.g., a sentence-encoder embedding of a per-pair rationale ("oil-sector earnings beat" → encoded with mpnet), retrieved by manual labeling or a small LLM. --- ### HIGH-5: AMP Scaler State Corruption in NaN-Guard **Where:** `pipeline.py:1363-1371`. **Evidence:** ```python if not torch.isfinite(loss): nan_skip_count += 1 optim.zero_grad(set_to_none=True) scaler.update() # <-- called without scaler.scale(loss).backward() ... continue ``` The recommended pattern in pytorch docs is: if you skip a step, you must NOT call `scaler.update()` because the scaler's growth tracker counts only valid steps. Calling `update()` after a skip prematurely grows the loss scale, making subsequent NaN events more likely. **Why it matters:** Once you've started accumulating NaN events (which the code itself notes happen — "50 events at a time"), this guard accelerates the problem rather than fixes it. Training stability is at risk. **Fix:** Remove the `scaler.update()` call inside the NaN branch. Pytorch's recommended pattern: ```python if not torch.isfinite(loss): optim.zero_grad(set_to_none=True) continue # Do NOT touch the scaler. ``` --- ### MEDIUM-1: Direction is 3-Class but Effectively 2-Class **Where:** `pipeline.py:1067` `'direction': int(np.sign(rret))`; then `pipeline.py:1313` `int(r['direction']) + 1` maps to CE labels {0, 1, 2}. **Evidence:** `np.sign(continuous_return)` is essentially never exactly 0. So the "flat" class is never observed. CrossEntropyLoss over a 3-class head with class 1 (flat) absent biases the head's logits to two-class behavior with a constant offset for the unused class. **Why it matters:** Wastes capacity; dir_acc can plateau at the larger of the up/down class proportions without learning anything. **Fix:** Either make direction 2-class (up/down) or define "flat" as `|return| < 0.5 * pre_vol`. Currently dir is conceptually 3-class but practically 2-class with a deadweight third class. --- ### MEDIUM-2: Symbol Coercion Table is Off-by-One and Sloppy **Where:** `pipeline.py:307-338`. **Evidence:** - Claim: "27-entry _SYMBOL_COERCIONS". - Actual: 26 entries (counted programmatically). - One identity: `'^MERV': '^MERV'` (line 311) — does nothing. - One circular pair: `'^TPX': '^TOPX'` (line 310) and `'^TOPX': '^TPX'` (line 337). Apply the coercion twice and you cycle. **Why it matters:** Symptomatic of unreviewed copy-paste edits to a shared mapping. The circular pair could be a latent bug if any code path applies coercion iteratively (it doesn't, but the door is open). **Fix:** Pick one canonical Tokyo Topix symbol and route both ^TPX and ^TOPX to it. Drop `'^MERV': '^MERV'` (or change to comment that it's an explicit no-op). Update claim to 27 actual entries. --- ### MEDIUM-3: No Test Files Exist **Where:** Whole project. **Evidence:** `find` for `test_*.py`, `*_test.py`, `conftest.py` returns nothing under `merged_news/`. **Why it matters:** Future regressions (especially in `compute_impact_label` and `HybridNewsImpactModel.forward`) will go undetected. The two CRITICAL bugs above would have been caught by 10 lines of pytest each. **Fix:** Add `tests/test_pipeline.py` with the recurrence-prevention tests listed in each CRITICAL finding. --- ### LOW-1: Deprecated APIs - `torch.cuda.amp.autocast` / `torch.cuda.amp.GradScaler` — deprecated since pytorch 2.0 in favor of `torch.amp.autocast('cuda', ...)`. Currently emits DeprecationWarning at runtime. - `datetime.datetime.utcnow()` (pipeline.py:46, 1484) — deprecated since python 3.12 in favor of `datetime.datetime.now(timezone.utc)`. Currently emits DeprecationWarning. Fixes are mechanical and one-line each. --- ## 4. Model Architecture Verification | Component | Declared | Actual (from code) | Compatible with `final_best.pt`? | |-----------|----------|--------------------|----------------------------------| | Text encoder | roberta-large 336M FROZEN | ✓ wired correctly when `freeze_text=True` | UNKNOWN (no ckpt access) | | Chart encoder | small Transformer over 128-step OHLCV | `ChartEncoder`: in=5, hidden=128, layers=4, nhead=8, out=text_dim. Pools with `mean(dim=1)`. | UNKNOWN | | Fusion | 8-head cross-attention 1024-dim | `nn.MultiheadAttention(1024, 8, batch_first=True)` — *but collapses to identity-on-value due to T=1* | UNKNOWN | | `head_sig` | Linear(1024, 2) | ✓ | UNKNOWN | | `head_dir` | Linear(1024, 3) | ✓ | UNKNOWN | | `head_mag` | Linear(1024, 1) | ✓ | UNKNOWN | | `head_reason` | Linear(1024, 1024) | ✓ — **but receives no loss signal**, weights remain at init | UNKNOWN | | Trainable params | 6.2M | ~6.5M expected (ChartEncoder ≈ 0.6M + fuse ≈ 4.2M + 4 heads ≈ 1.1M). Run script H2 to verify. | UNKNOWN | The architecture *as written* matches the claims at the layer-construction level. The collapse bug is a *forward-pass semantics* issue, not a layer-shape issue. Hence `final_best.pt` likely loads strict=True without errors; the silent failure is in the data flow, not the weight shapes. --- ## 5. Data & Label Audit ### News distribution | Property | Value | |----------|-------| | Total rows | 9,646,340 | | Columns | 4 (title, summary, date, source) — all `large_string` | | Date range | 1999-01-01 → 2026-05-03 | | Source: The Guardian | 2,525,498 (26.2%) | | Source: BBC | 7,120,842 (73.8%) | | Date-null rows | 3,504,560 (36.3%) | | Date-null rows in row-group 0 | 29,467 = exactly the BBC count in rg0 (suspicious) | | First-rg duplicate titles | 152,920 / 1,048,576 = 14.6% | | First-rg duplicate (title, date) | 2,052 (0.2%) | ### Effective training corpus (after losses): ``` 9.6M raw news - 3.5M null-date (-36.3%) → 6.1M - news that mentions no symbol → unknown - news mentioning non-crypto symbols → labels rejected by 24h/daily issue ≈ "small fraction of crypto news, before deduplication" ``` ### Chart data (5 rescued, only sample available) | File | Rows | Date range | Resolution | OHLCV invariants | |------|------|-----------|------------|--------------------| | ETHUSD_X | 3,107 | 2017-11 → 2026-05 | daily | clean | | FBHS | 3,107 | 2014-01 → 2026-05 | daily | clean | | URA_F | 3,107 | 2014-01 → 2026-05 | daily | 1 minor low>close violation | | XAG_F | 3,107 | 2014-01 → 2026-05 | daily | 219 zero-volume bars | | idx_IBOV | 3,064 | 2014-01 → 2026-05 | daily | 30 zero-volume bars | All five are daily, contradicting the directory name `charts_4h/`. Zero-volume bars for futures (XAG_F) and indices (IBOV) are expected (holidays, low-liquidity sessions) and not bugs — but `pre_vol` computed across these bars could include zero-return periods that artificially shrink volatility. ### Label trace: what happens to a typical equity news event? 1. News at 2024-03-15 14:30 UTC: "Apple beats earnings". 2. `attribute_news_to_symbols` matches "Apple" → symbol AAPL. 3. `build_training_pairs` reads `charts_4h/AAPL.parquet` — daily bars. 4. `compute_impact_label` finds: - pre window (2024-03-14 14:30 → 2024-03-15 14:30): up to 1 daily bar (the 2024-03-15 EOD bar may be tomorrow's close, hence pre window may have 0-1 bars depending on UTC timing of EOD). - `pre['close'].pct_change().std()` over 0-1 values → NaN. - `if pd.isna(pre_vol): return {}` → labeled rejected. 5. Pair contributes nothing to training. AAPL effectively contributes nothing. ### Label trace: a crypto news event 1. News at 2024-03-15 14:30 UTC: "Bitcoin surges past $70k". 2. Matches symbol BTC/USDT (via crypto-augmented aliases). 3. Reads ccxt-fetched 4h bars. 4. pre window 24h = 6 bars; pct_change → 5 valid returns; std OK. 5. `excess = |48h return| / 4h-bar-std`. 6. With BTC's typical 4h vol of ~1-2% and 48h returns often in the 2-5% range, `excess ≈ 1.5-5`. Many will cross threshold 3.0. **Significance rate ~50% → majority class**. This trace matches the observed sig_acc plateau of 0.557 — a chart-only model on a crypto-heavy corpus with sig threshold inflated by dimensional mismatch. --- ## 6. Metric Audit ### What does the loss decompose to? `val_loss = ce(sig) + ce(dir) + 0.5 * smoothl1(mag)` Baseline reference values: | Baseline | ce(sig) | ce(dir) | 0.5·smoothl1(mag) | total | |----------|---------|---------|--------------------|-------| | Random (uniform) | ln 2 ≈ 0.693 | ln 3 ≈ 1.098 | ~0.5 * E[|mag|] (depends on dist) | ≈ 2.0-2.5 | | Majority class sig (55.7%) | ≈ 0.685 | — | — | — | | Majority class dir (probably ~50%) | — | ≈ 0.693 | — | — | | Chart-only model (observed) | ? | ? | ? | 2.495 | The total of 2.495 is consistent with: majority-class sig (~0.69) + slightly better than majority dir (~1.0) + a partially trained mag head (~0.8 × 0.5 = 0.4), summed gives ~2.09 — under 2.495. So the model is probably worse than that decomposition would suggest, or the mag loss dominates. Without per-component val losses in the report, this can't be pinned down. ### v6 vs v4: is 0.0004 meaningful? No. Without seed-controlled triplicates: - Two-sample t-test requires n ≥ 2 per condition with std estimate. - The session ran each config once. Sample size = 1. - Run-to-run std of converged validation loss on this kind of model is typically 0.01-0.05. - Δ = 0.0004 is two orders of magnitude below the noise floor. **Verdict: "v6 winner" is a coin flip with extra steps.** ### Suggested baselines that were NOT run These would distinguish "model learns from news" from "model learns prior": - **Chart-only ablation** — drop text encoder entirely (or pass random text); measure val loss. If equal to current v6, the model isn't using text. (Given C16, this baseline is *guaranteed* equal.) - **Text-only ablation** — drop chart encoder; measure val loss. If much higher, news semantics are useful; if equal, news isn't being used. - **Permuted-label baseline** — shuffle labels within a shard before training. If val loss drops similarly, the model is learning train/val correlations unrelated to labels. - **Permuted-text baseline** — pair each chart with a randomly-assigned news text from the same shard. If val loss is unchanged from baseline, text is unused (which we know is the case). --- ## 7. Inference Audit ### What's verifiable from local code - The inference snippet in the prompt is *almost* correct, but: - `model.load_state_dict(state['model'] if 'model' in state else state)` — in `train_one_shard` (line 1437-1443) the checkpoint key is `'state'`, not `'model'`. In `merge_shards_and_test` (line 1475) it's also `'state'`. So the snippet should be `state['state'] if 'state' in state else state`. - `state.get('model', state)` is *forgiving* and won't error — it just falls through. But the documented-correct line would use 'state'. ### Hypothetical robustness failure modes Given the cross-attention collapse, all "model robustness" tests degenerate to "chart robustness". Specifically: - **Same chart, different news** → output is bit-identical (predicted). - **Empty/garbled news** → output identical to legitimate news (predicted). - **Adversarial news ("guaranteed crash")** → output unchanged (predicted). - **NaN chart input** → likely outputs NaN; no defensive handling in forward. - **Sub-128-step chart** → in `train_one_shard._DS.__getitem__` lines 1298-1300, the chart is zero-pre-padded if shorter — but at inference time the caller is responsible for windowing. No safety net. ### Shape contract The forward returns shapes consistent with the documented contract: - `sig_logits`: (B, 2) ✓ - `dir_logits`: (B, 3) ✓ - `mag`: (B,) — squeezed from (B,1) ✓ - `reason_emb`: (B, 1024) for roberta-large ✓ The shape contract is OK; the semantic contract (model uses news to predict) is not. --- ## 8. Code Bug Hunt — Function-Level Summary | Function | LOC | Status | Notes | |----------|-----|--------|-------| | `normalize_date_column` | 5 | OK | Simple wrapper | | `merge_all_news` | ~40 | OK | Standard concat + dedupe | | `build_symbol_pool` | ~170 | NOT REVIEWED — UNKNOWN | Large; should be reviewed for pool composition | | `fetch_yfinance_daily` | ~70 | OK / FAIL-RESOLUTION | Function works; resolution implication is the real bug | | `fetch_stooq_daily` | ~45 | OK | Fallback path | | `fetch_crypto_universe_from_exchanges` | ~85 | NOT REVIEWED | | | `fetch_ccxt_4h` | ~50 | OK | Correctly requests 4h | | `fetch_all_charts` | ~60 | PARTIAL | Saves daily and 4h into one dir; loud rename suggested | | `build_symbol_aliases` | ~120 | NOT REVIEWED — UNKNOWN | Long; should be audited for ticker collisions (e.g. "X", "META", "T") | | `attribute_news_to_symbols` | ~70 | OK | Vectorized findall; correct algorithm. Common-word collision risk depends on `build_symbol_aliases` content. | | `compute_impact_label` | ~40 | FAIL | C12, C14, C15 above | | `get_model_classes` | ~60 | FAIL (C16) | Cross-attention collapse | | `run_data_prep` | ~6 | OK | Orchestrator | | `run_feature_prep` | ~50 | OK | Shard-then-augment pipeline; `iloc[shard::num_shards]` interleaves correctly | | `_safe_symbol_filename` | 2 | OK | | | `build_training_pairs` | ~25 | PARTIAL | Loops one symbol at a time; could be vectorized. Inherits compute_impact_label issues | | `train_one_shard` | ~225 | PARTIAL | C30 (temporal split) OK, C35 (no seed) FAIL, C37 (scaler bug) FAIL, no head_reason supervision | | `merge_shards_and_test` | ~40 | FAIL (C26) | Reports pre-merge val, not merged-model test | | `multi_iterate` | ~110 | PARTIAL | Uses `<` for best-tracking with no statistical confidence (line 1569) | | `iterate_v2`, `train_remaining_shards_and_merge` | — | NOT REVIEWED | Lower priority for v6 model audit | --- ## 9. Production Readiness | Dimension | Status | Notes | |-----------|--------|-------| | Checkpoint integrity | UNKNOWN | No Drive access | | Architecture matches spec | FAIL | C16 | | Reproducibility | FAIL | C35 | | Tested code | FAIL | C38 | | Logging / metrics | PARTIAL | Per-step JSONL logging exists, but per-component val loss not logged | | Dependency pinning | UNKNOWN | No requirements.txt / pyproject in scope | | Model registry | NONE | Drive-only artifacts | | Calibration | NONE | No calibration metrics computed | | Input validation | NONE | No defensive code at inference boundary | | "Reason" output usability | FAIL | C36 — embedding, not explanation | | Ops state at session close | PARTIAL | Cron disabled (good), tunnel still up (risk) | | Security: `torch.load(pt)` | RISK | Pickle deserialization; OK if artifact is trusted | | Data freshness | PARTIAL | News ends 2026-05-03; close-to-date OK | **Verdict: not deployable as a "news-impact AI". Deployable, with significant caveats, as a chart-only crypto move classifier — which is not what's advertised.** --- ## 10. Recommended Next Actions ### Do immediately (under 1 day) 1. Fix the cross-attention collapse (CRITICAL-1). Either Option A (concat-then-project) or Option B (chart-as-sequence). Add the "text-affects-output" pytest. **This is the single highest-ROI fix.** 2. Wire `head_reason` either to a real target or remove it (HIGH-4). 3. Fix the AMP NaN-guard (HIGH-5). 4. Set seeds at the top of every runner; tag the config with seed in the manifest (HIGH-3). ### Before v7 training 5. Audit and re-derive `compute_impact_label`: - Pick a single bar resolution (4h or daily) per symbol and *fail loudly* if mismatched. - Fix the vol-window ↔ return-window dimensional mismatch (HIGH-1). - Either drop the "flat" direction class or define it as `|return| < ε·vol` (MEDIUM-1). 6. Add a held-out global test set; have `merge_shards_and_test` evaluate the merged model on it (CRITICAL-3). 7. Re-parse BBC news dates (HIGH-2). Confirm ingestion isn't silently dropping timestamps. 8. Add the 4 baselines from §6: chart-only, text-only, permuted-label, permuted-text. Report all 5 numbers in the iteration scoreboard. 9. Run each config 3 times with seeds {42, 1337, 7} and report mean ± std. ### Before production deploy 10. Add `tests/test_pipeline.py` with at minimum: - `test_text_input_affects_output` (CRITICAL-1) - `test_compute_impact_label_handles_daily_bars` (CRITICAL-2) - `test_merge_then_evaluate_matches_report` (CRITICAL-3) - `test_significance_threshold_calibrated` (HIGH-1) 11. Pin dependencies (`requirements.txt` or `uv.lock`). 12. Verify `final_best.pt` loads strict=True; run `audit_inference_skeleton.py`. 13. Decide what `reason_emb` actually does in the product. Either build a retrieval bank, train a small explanation decoder, or replace with feature attribution. 14. Add monitoring: log inference latency, input shape, output distribution. ### Nice to have 15. Replace `datetime.utcnow()` and `torch.cuda.amp.*` with their non-deprecated equivalents (LOW-1). 16. Clean up `_SYMBOL_COERCIONS` (MEDIUM-2). 17. Vectorize `build_training_pairs` (it currently iterates symbol-by-symbol, row-by-row). --- ## 11. Hypothesis Refutation — Acımasız Mod Sonuçları The user asked me to test 10 negative hypotheses. Status for each: | # | Hypothesis | Status | Confidence | Evidence | |---|-----------|--------|------------|----------| | 1 | Model only learned class prior, not real signal | **CONFIRMED** | 1.0 | C16 forces this; sig_acc = majority class | | 2 | Validation split contains leakage | **REFUTED** | 0.9 | Per-shard temporal split is clean; cross-shard temporal overlap is mitigated by per-shard val | | 3 | Label function is noisy/misaligned | **CONFIRMED** | 1.0 | C12 (daily/4h mismatch), C14 (vol dim mismatch), C15 (flat class) | | 4 | reason head produces no explanation | **CONFIRMED** | 1.0 | C36 — 1024-dim float vector, no decoder, no training target | | 5 | v6 winner claim is statistically meaningless | **CONFIRMED** | 1.0 | C27, C35 — no seed, n=1 per config, Δ < noise band | | 6 | final_best.pt config doesn't match report | UNKNOWN | — | Cannot verify w/o Drive access | | 7 | Chart input contains future leakage | PARTIAL | 0.7 | Daily-bar timestamps may include a "today's close" that postdates news arrival; window logic is `Date <= t`. For intraday news on daily data, you'd grab same-day EOD which is future. **Suspected leakage for non-crypto.** | | 8 | News-symbol attribution has high false positives | UNKNOWN | — | Depends on `build_symbol_aliases` content; not audited in this pass | | 9 | Loss plateau is architecture/label, not capacity | **CONFIRMED** | 0.95 | C16 + C12 + C14 jointly explain the plateau | | 10 | Production inference snippet is fragile | **CONFIRMED** | 0.9 | Uses wrong checkpoint key (`'model'` instead of `'state'`); falls through by luck. No input validation. | --- ## 12. Answer to "Bu modele güvenebilir miyim?" 1. **`final_best.pt` gerçekten yüklenip inference yapıyor mu?** UNKNOWN (Drive yok). Mimari shape uyumu var → muhtemelen yükleniyor. `audit_inference_skeleton.py` ile doğrulayın. 2. **Mimari iddia ile checkpoint uyumlu mu?** Çok büyük olasılıkla evet (shape düzeyinde). Ama mimari SEMANTİK olarak bozuk (C16). 3. **Raporlanan metrikler doğru ve anlamlı mı?** Hayır. 2.4951, post-merge test loss değil; pre-merge per-shard val ortalaması. Bu farkı düzeltmeden iddia "winner" olamaz. 4. **v6 gerçekten daha iyi mi yoksa noise mu?** Noise. Δ=0.0004, seed yok, sample size 1. 5. **Data/label noise gerçek bottleneck mi?** Kısmen. Label noise gerçek (C12, C14, C15) ama asıl bottleneck mimari (C16). Mimari düzeltilmeden label'ları düzeltmek hiçbir şey çözmez. 6. **Major bug/leakage var mı?** Evet, üç kritik bug: - Cross-attn collapse - Daily/4h resolution silent rejection - Pre-merge metric reporting 7. **Production deploy güvenli mi?** Hayır. 8. **v7'den önce mutlaka düzeltilmesi gereken 5 şey:** 1. Cross-attention fusion (CRITICAL-1) 2. Chart resolution mismatch (CRITICAL-2) ve label window 3. Post-merge evaluation (CRITICAL-3) 4. Seed control + triplicate runs 5. `head_reason`'u ya kaldır ya da gerçekten eğit --- *End of audit report. Generated by Claude on 2026-05-13 from local artifacts. Drive-bound claims marked UNKNOWN; run `audit/audit_inference_skeleton.py` on Colab to convert UNKNOWNs to PASS/FAIL.*