--- pretty_name: AutoDataBench Knowledge Injection Resources tags: - autodatabench - knowledge-injection - question-answering --- # AutoDataBench Knowledge Injection Resources Public resources for the knowledge-injection task in [AutoDataBench](https://github.com/AutoDataBench/AutoDataBench). See the [paper](https://arxiv.org/abs/2609.40097) for the benchmark setting. ## Contents ```text data/knowledge_injection_v1/context_pool.jsonl data/knowledge_injection_v1/sources.jsonl models/talkie-1930-13b-it-vllm/ models/Qwen3-4B-Instruct-2507/ models/Qwen3-Embedding-0.6B/ ``` - `sources.jsonl` contains 1,000 benchmark-relevant post-1930 Wikipedia summaries. - `context_pool.jsonl` contains the full 542,970-row post-1930 retrieval corpus. The 1,000 target sources are included in this pool. Neither file contains evaluation questions, answer choices, or labels. | Model | Role | Original model | | --- | --- | --- | | talkie-1930-13b-it-vllm | Fixed knowledge-injection base model | [awilliamson/talkie-1930-13b-it-vllm](https://huggingface.co/awilliamson/talkie-1930-13b-it-vllm) | | Qwen3-4B-Instruct-2507 | Agent-callable generation model | [Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) | | Qwen3-Embedding-0.6B | Agent-callable embedding model | [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) | ## Evaluation data The 1,000 novel-knowledge probes and 4,400 retention probes are evaluator-only and are intentionally excluded. Keep those splits outside the agent sandbox when running the benchmark. ## Use with AutoDataBench Copy or symlink `data/` and `models/` into the AutoDataBench repository. The paths already match the default task configuration. Point the generation and embedding servers at the local auxiliary-model directories if needed. `MANIFEST.sha256` contains checksums for every distributed file. Model and dataset components retain their upstream licenses. Consult the model cards and source datasets before redistribution or commercial use. ## Citation If you use these resources, please cite: ```bibtex @misc{yuan2026autodatabench, title = {AutoDataBench: A Data-centric Testbed for Accelerating Auto Research}, author = {Ruifeng Yuan and Yizhi Li and Yaxin Du and Fengyu Cai and Yiqi Liu and Hou Pong Chan and Chenghua Lin and Yun Chen and Jian Yang and Bryan Dai and Pinyan Lu and Chenghao Xiao}, year = {2026}, eprint = {2609.40097}, archivePrefix = {arXiv}, primaryClass = {cs.CL}, url = {https://arxiv.org/abs/2609.40097} } ```