mabera commited on
Commit
6d4b751
·
verified ·
1 Parent(s): 4f103f8

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +71 -0
README.md CHANGED
@@ -1,3 +1,74 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ language:
4
+ - en
5
+ tags:
6
+ - math-code
7
+ - mathematics
8
+ - verification
9
+ - closed-label
10
+ - autoscientist
11
+ base_model: mistralai/Mixtral-8x7B-Instruct-v0.1
12
  ---
13
+
14
+ # Math Solution Verification Classifier
15
+
16
+ **Author:** Hussein Adeiza (mabera)
17
+ **Role:** Licensed Environmental Health Officer, Abuja Nigeria
18
+ **Base Model:** Mixtral 8x7B
19
+ **Fine-tuned with:** AutoScientist by Adaption Labs
20
+
21
+ ## Model Description
22
+ A LoRA adapter fine-tuned to verify whether a proposed answer to a real
23
+ competition math problem is correct, classifying it as Correct or
24
+ Incorrect. This is a genuinely global-scope submission (not Nigeria-
25
+ specific), addressing a universal AI capability question: can a model
26
+ reliably grade mathematical correctness?
27
+
28
+ ## Training Data
29
+ - Source: MATH dataset (Hendrycks et al., NeurIPS 2021), accessed via
30
+ a properly-cited GitHub derivative (rasbt/math_full_minus_math500),
31
+ downloaded directly, 12,000 real competition problems
32
+ - Dataset: 20 rows (10 real problems x 2 answer variants each: one
33
+ correct, one deliberately perturbed incorrect answer, disclosed as
34
+ a constructed perturbation, not a real student error)
35
+ - Kaggle: https://www.kaggle.com/datasets/yunusahusseinadeiza/math-solution-verification-classifier
36
+
37
+ ## Important Note: Column Selection Correction
38
+ During training setup, the platform defaulted to training on
39
+ "Enhanced completion" text, which had drifted away from the closed-
40
+ label Correct/Incorrect structure into generic step-by-step tutoring
41
+ language, losing the classification task entirely. This was caught
42
+ and manually corrected by selecting "Original completion" instead
43
+ before training. Worth flagging for other builders working on
44
+ closed-label tasks: check which completion column is actually
45
+ selected before training, since the platform default may not be
46
+ the one you expect.
47
+
48
+ ## Training Metrics
49
+ - **Win rate (on dataset): 76% adapted vs 24% base model**
50
+ - Base model: mistralai/Mixtral-8x7B-Instruct-v0.1
51
+ - Method: LoRA (confirmed via training config), no recipe modifications
52
+ - Dataset quality: 7.0 → 9.0 (+28.6% relative improvement, **Grade A**)
53
+ - Percentile: 33.0
54
+ - Domain classification: Math (100%), a clean, accurate match
55
+
56
+ ## Verification
57
+ All 20 rows independently, programmatically verified before training:
58
+ every Correct/Incorrect classification checked against the real
59
+ ground-truth answer looked up directly in the raw MATH dataset source
60
+ file. 20/20 pass rate, demonstrated live in the accompanying Kaggle
61
+ notebook.
62
+
63
+ ## Why This Result Matters
64
+ This is the highest quality grade (A) and highest quality score
65
+ improvement (+28.6%) in this author's 18-submission AutoScientist
66
+ portfolio, achieved using the same disciplined closed-label methodology
67
+ (no recipe modifications, deterministic ground truth, independent
68
+ verification) established across the AMR and Loan Classifier
69
+ submissions, applied here to a genuinely global rather than
70
+ Nigeria-specific problem.
71
+
72
+ ## Credits
73
+ Powered by Adaptive Data — Adaption Labs
74
+ AutoScientist Challenge 2026, Part 2 — Math & Code Category