Qwen 3.5 9B NL2SQL LoRA

This repository contains the final LoRA adapter for the Qwen 3.5 9B model line used in a master's thesis project on local large language models for NL2SQL generation.

The adapter generates SQLite queries from natural-language questions and a relational database schema. It must be loaded together with the corresponding base model.

Base model

  • Model: Qwen/Qwen3.5-9B-Base
  • Revision: 68c46c4b3498877f3ef123c856ecfde50c39f404

Adapter

  • Method: LoRA supervised fine-tuning
  • LoRA rank: 8
  • LoRA alpha: 16
  • LoRA dropout: 0.05
  • Target modules: all suitable linear modules
  • Quantization during training: none
  • Maximum training sequence length: 2,048 tokens
  • Best checkpoint: checkpoint-502
  • The published root adapter corresponds to the selected best checkpoint.

SHA-256 of adapter_model.safetensors:

dddf120df0703be5b9106ba17a628f2a9664e6ab5d1cc3ec1311c0a4a2b000f0

Training configuration

  • Training examples: 25,000
  • Spider Train examples: 6,960
  • SQL Create Context examples: 18,040
  • Validation set: MixedVal2500-v2
  • Validation examples: 2,500
  • Learning rate: 1e-4
  • Scheduler: constant
  • Train batch size: 2
  • Gradient accumulation steps: 4
  • Effective batch size: 8
  • Seed: 42
  • Maximum epochs: 5
  • Early stopping patience: 2
  • Early stopping threshold: 0.001
  • Precision: FP16
  • Gradient checkpointing: enabled
  • Attention implementation: FlashAttention 2

Spider Dev was not used for training, validation, early stopping, or checkpoint selection.

Evaluation

The final adapter was evaluated on all 1,032 Spider Dev cases.

Zero-shot evaluation:

  • Execution Match Accuracy: 74.52% (769/1,032)
  • Execution Success Rate: 97.00% (1,001/1,032)
  • Maximum input length: 2,048 tokens
  • Maximum generated tokens: 256

These values correspond to the authoritative evaluation run documented in the accompanying GitHub repository.

Intended prompt behavior

The model is instructed to return only a valid SQLite query:

  • no explanation
  • no Markdown
  • no comments
  • no unnecessary tables or columns
  • only SELECT or WITH queries
  • output terminated with a semicolon

The exact prompt construction and evaluation pipeline are documented in the project repository.

Loading

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_model_id = "Qwen/Qwen3.5-9B-Base"
adapter_id = "mehmet1899/qwen35-9b-nl2sql-lora"
adapter_revision = "e136b9c25ede3ee82210875d0db774089509b676"

tokenizer = AutoTokenizer.from_pretrained(
    base_model_id,
    revision="68c46c4b3498877f3ef123c856ecfde50c39f404",
)

model = AutoModelForCausalLM.from_pretrained(
    base_model_id,
    revision="68c46c4b3498877f3ef123c856ecfde50c39f404",
    torch_dtype=torch.float16,
    device_map="auto",
)

model = PeftModel.from_pretrained(
    model,
    adapter_id,
    revision=adapter_revision,
)
model.eval()

Reproducibility

Code, training and evaluation configurations, environment information, run manifests, and result summaries are available at:

https://github.com/md181099/nl2sql-masterthesis

The files training_metadata.json, training_history.csv, and training_history.jsonl provide additional training provenance.

Limitations

  • The adapter was evaluated primarily on the Spider benchmark and SQLite databases.
  • Performance on other database systems or unseen schema conventions is not guaranteed.
  • Execution Match depends on the database contents and the execution-based evaluation procedure.
  • The model may still generate invalid, incomplete, or semantically incorrect SQL.
  • The adapter should not be used to execute unrestricted queries against production databases without validation and access controls.
Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mehmet1899/qwen35-9b-nl2sql-lora

Adapter
(13)
this model