--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation base_model: - Qwen/Qwen3-8B tags: - code - software-engineering - agent --- # FIM-8B Inference on SWE-Bench Verified This guide describes how to run the **FIM-8B** checkpoint on SWE-Bench Verified (and Lite). Unlike FIM-7B and FIM-14B (R2E-Gym scaffold), this model is post-trained on SWE-Lego trajectories and is evaluated with the SWE-Lego setup: OpenHands `CodeActAgent` for inference and the official SWE-bench harness for scoring. ## Model Local path: `models/FIM-8B/` (checkpoints are gitignored; do not commit them). - Base model: `Qwen/Qwen3-8B` - FIM mid-training: FIM v2/v3 data - Post-training: SFT on SWE-Lego trajectories ## 1. Serve the model with vLLM ```bash CUDA_VISIBLE_DEVICES=0 \ python -m vllm.entrypoints.openai.api_server \ --model models/FIM-8B \ --served-model-name FIM-8B \ --host 127.0.0.1 \ --port 8400 \ --tensor-parallel-size 1 \ --max-model-len 163840 \ --max-num-seqs 16 \ --gpu-memory-utilization 0.9 \ > vllm_fim8b.log 2>&1 & ``` The checkpoint ships `max_position_embeddings: 163840` and its own chat template, so no rope or template overrides are needed. Wait until the server is up (model load takes ~1 minute): ```bash curl -s http://127.0.0.1:8400/v1/models ``` ## 2. Run the agent on SWE-Bench Verified Inference uses OpenHands 0.53.0. Define the LLM in `config.toml`: ```toml [llm.eval_fim] model = "openai/FIM-8B" base_url = "http://127.0.0.1:8400/v1" api_key = "EMPTY" temperature = 0.0 max_input_tokens = 147456 max_output_tokens = 16384 native_tool_calling = false ``` From the OpenHands checkout: ```bash env USE_HINT_TEXT=false \ INSTRUCTION_TEMPLATE_NAME=swe_default.j2 \ ENABLE_PLAN_MODE=false \ ADD_IN_CONTEXT_LEARNING_EXAMPLE=false \ poetry run python evaluation/benchmarks/swe_bench/run_infer.py \ --config-file config.toml \ --agent-cls CodeActAgent \ --llm-config llm.eval_fim \ --max-iterations 100 \ --eval-num-workers 1 \ --eval-output-dir ./eval_out \ --dataset princeton-nlp/SWE-bench_Verified \ --split test \ --mode swe ``` For SWE-Bench Lite, use `--dataset princeton-nlp/SWE-bench_Lite`. ## 3. Score with the SWE-bench harness Convert the OpenHands `output.jsonl` to a predictions file with `evaluation/benchmarks/swe_bench/scripts/eval/convert_oh_output_to_swe_json.py`, then evaluate it with the official SWE-bench harness (`python -m swebench.harness.run_evaluation`). The reported score is `resolved_instances / total_instances`.