DIY Open Pangram - Part 1
Data
I asked ChatGPT codex to launch a batch of luna agents to go find some good datasets for this task, roughly based on the data mix percentages from the Pangram 4 paper.
We ended up with:
- PMC + ACL + EditLens combined: 4,404 rows (Qwen3 and combined baselines)
- PMC + ACL papers: 1,542 rows (paper baselines)
- EditLens general text: 37,568 available; baselines used 20,000 rows
- PMC-only abstracts: 616 rows (PMC-only baselines)
Baselines
I wanted to have a good set of baselines to compare against, and ideally beat. So I calibrated and evaluated a few baseline methods/models on our evaluation dataset:
- Character 3-5-gram TF-IDF + logistic regression
- Word 1-2-gram TF-IDF + logistic regression
- MiniLM embeddings + logistic regression
- EditLens RoBERTa-large
- EditLens Llama-3.2-3B
- Load-bearing PR vocabulary model
First training run
I left our model to do its first training run overnight. I forgot to report to wandb, so generated a loss chart after the fact:
Apparently the best checkpoint was from half way through training, this is a little concerning. So i think I want to tune the training hyperparameters next. I'd like to have some sense of training constantly improving the model, and also decrease variance across seeds.
Evaluation
The evaluation of the model seemed a little bit mediocre at first, getting reasonable AI recall rate but bad false positives, around 4%. I asked it to change the threshold and then it reported suspiciously high results:

Looking at the ROC charts also concernes me

So I think there is something wrong that I need to figure out about these "too good to be true" results before continuing.
It seems like there was miscommunication between myself and the AI - I said that part of the goal of this work was to do well on papers and it completely fixated on papers and only trained and evaluated to that. I'm now trying to widen the dataset significantly over a broad range of document types where we have some labels for whether it's human or AI text.