SourceCodeAuthorCheck-SLM-10M
A ~10 million parameter Small Language Model (SLM) Transformer designed to detect whether a Python source code file was written by a human or generated by an AI model.
Live Demo
Test the model directly in your browser without writing code: SourceCodeAuthorCheck Web UI
π¬ Training Pipeline & Techniques
This model was built from scratch using a custom PyTorch TransformerEncoder architecture (~9.6M parameters). The training pipeline utilized several specific techniques to ensure accurate binary classification and optimal hardware utilization.
1. Temporal Data Separation
To create a stark contrast between human and AI coding paradigms, the dataset relies on temporal splitting:
- Human Baseline (Class 0): Python source code extracted from GitHub repositories created in Q3 2017 and prior, guaranteeing the code predates modern generative AI.
- GenAI Baseline (Class 1): Synthetic datasets structured to mimic the exact architectural paradigms, repetitive docstrings, and token distributions typical of models operating in Q3 2026.
2. Hardware Optimization (NVIDIA DGX Spark)
The model was trained natively on an NVIDIA DGX Spark (Grace Blackwell architecture).
- Automatic Mixed Precision (AMP): We utilized PyTorch's
torch.autocasttargetingbfloat16. This leverages Blackwell's 5th-generation Tensor Cores, accelerating matrix multiplications while maintaining numerical stability during backpropagation. - Gradient Scaling: Paired with AMP,
torch.amp.GradScalerwas used to prevent underflow errors during the transition between FP32 and BF16 formats.
3. Optimization & Loss
- Loss Function:
BCEWithLogitsLoss. This combines a Sigmoid layer and Binary Cross Entropy Loss in a single class, providing better numerical stability than applying Sigmoid followed by standard BCELoss. - Optimizer:
AdamW(Adam with Weight Decay) to enhance generalization and prevent overfitting on the synthetic AI subsets.
π» Usage: The Inference Script
The easiest way to use this model locally is via the standalone inference.py script included in this repository. It includes the required architecture class and handles downloading the weights automatically.
1. Download the script You can download the script directly from the files tab: inference.py
2. Run against any Python file Pass the path of the file you want to analyze directly to the script:
python inference.py my_script.py
Example Output:
--- Testing File: my_script.py ---
Verdict: Human Written (AI Probability: 37.89%)
Preview: def process_data(items):...
π Programmatic Usage
If you want to integrate the model directly into your own Python applications, you must define the architecture class before loading the weights.
import torch
import torch.nn as nn
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
# 1. Define the Architecture
class SourceCodeAuthorCheck(nn.Module):
def __init__(self, vocab_size=50257, d_model=128, nhead=8, num_layers=4, dim_feedforward=512):
super().__init__()
self.embedding = nn.Embedding(vocab_size, d_model)
self.pos_encoder = nn.Parameter(torch.zeros(1, 1024, d_model))
encoder_layers = nn.TransformerEncoderLayer(
d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, batch_first=True
)
self.transformer = nn.TransformerEncoder(encoder_layers, num_layers=num_layers)
self.fc = nn.Linear(d_model, 1)
def forward(self, input_ids, attention_mask):
seq_len = input_ids.size(1)
x = self.embedding(input_ids) + self.pos_encoder[:, :seq_len, :]
src_key_padding_mask = ~attention_mask.bool()
x = self.transformer(x, src_key_padding_mask=src_key_padding_mask)
mask_expanded = attention_mask.unsqueeze(-1).float()
sum_embeddings = torch.sum(x * mask_expanded, 1)
sum_mask = torch.clamp(mask_expanded.sum(1), min=1e-9)
pooled = sum_embeddings / sum_mask
return self.fc(pooled)
# 2. Load Tokenizer and Model Weights
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
model = SourceCodeAuthorCheck().to(device)
model_path = hf_hub_download(repo_id="assix-research/SourceCodeAuthorCheck-SLM-10M", filename="source_code_classifier.pth")
model.load_state_dict(torch.load(model_path, map_location=device, weights_only=True))
model.eval()
# 3. Analyze Code Snippet
code_snippet = "print('Hello World')"
inputs = tokenizer(
code_snippet, return_tensors="pt", truncation=True, padding="max_length", max_length=1024
).to(device)
with torch.no_grad():
logits = model(inputs['input_ids'], inputs['attention_mask'])
prob = torch.sigmoid(logits).item()
print(f"AI Probability: {prob:.1%}")