SourceCodeAuthorCheck-SLM-10M

Open in Spaces

A ~10 million parameter Small Language Model (SLM) Transformer designed to detect whether a Python source code file was written by a human or generated by an AI model.

Live Demo

Test the model directly in your browser without writing code: SourceCodeAuthorCheck Web UI


πŸ”¬ Training Pipeline & Techniques

This model was built from scratch using a custom PyTorch TransformerEncoder architecture (~9.6M parameters). The training pipeline utilized several specific techniques to ensure accurate binary classification and optimal hardware utilization.

1. Temporal Data Separation

To create a stark contrast between human and AI coding paradigms, the dataset relies on temporal splitting:

  • Human Baseline (Class 0): Python source code extracted from GitHub repositories created in Q3 2017 and prior, guaranteeing the code predates modern generative AI.
  • GenAI Baseline (Class 1): Synthetic datasets structured to mimic the exact architectural paradigms, repetitive docstrings, and token distributions typical of models operating in Q3 2026.

2. Hardware Optimization (NVIDIA DGX Spark)

The model was trained natively on an NVIDIA DGX Spark (Grace Blackwell architecture).

  • Automatic Mixed Precision (AMP): We utilized PyTorch's torch.autocast targeting bfloat16. This leverages Blackwell's 5th-generation Tensor Cores, accelerating matrix multiplications while maintaining numerical stability during backpropagation.
  • Gradient Scaling: Paired with AMP, torch.amp.GradScaler was used to prevent underflow errors during the transition between FP32 and BF16 formats.

3. Optimization & Loss

  • Loss Function: BCEWithLogitsLoss. This combines a Sigmoid layer and Binary Cross Entropy Loss in a single class, providing better numerical stability than applying Sigmoid followed by standard BCELoss.
  • Optimizer: AdamW (Adam with Weight Decay) to enhance generalization and prevent overfitting on the synthetic AI subsets.

πŸ’» Usage: The Inference Script

The easiest way to use this model locally is via the standalone inference.py script included in this repository. It includes the required architecture class and handles downloading the weights automatically.

1. Download the script You can download the script directly from the files tab: inference.py

2. Run against any Python file Pass the path of the file you want to analyze directly to the script:

python inference.py my_script.py

Example Output:

--- Testing File: my_script.py ---
Verdict: Human Written (AI Probability: 37.89%)
Preview: def process_data(items):...

πŸ›  Programmatic Usage

If you want to integrate the model directly into your own Python applications, you must define the architecture class before loading the weights.

import torch
import torch.nn as nn
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download

# 1. Define the Architecture
class SourceCodeAuthorCheck(nn.Module):
    def __init__(self, vocab_size=50257, d_model=128, nhead=8, num_layers=4, dim_feedforward=512):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, d_model)
        self.pos_encoder = nn.Parameter(torch.zeros(1, 1024, d_model))
        
        encoder_layers = nn.TransformerEncoderLayer(
            d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, batch_first=True
        )
        self.transformer = nn.TransformerEncoder(encoder_layers, num_layers=num_layers)
        self.fc = nn.Linear(d_model, 1)

    def forward(self, input_ids, attention_mask):
        seq_len = input_ids.size(1)
        x = self.embedding(input_ids) + self.pos_encoder[:, :seq_len, :]
        
        src_key_padding_mask = ~attention_mask.bool()
        x = self.transformer(x, src_key_padding_mask=src_key_padding_mask)
        
        mask_expanded = attention_mask.unsqueeze(-1).float()
        sum_embeddings = torch.sum(x * mask_expanded, 1)
        sum_mask = torch.clamp(mask_expanded.sum(1), min=1e-9)
        pooled = sum_embeddings / sum_mask
        
        return self.fc(pooled)

# 2. Load Tokenizer and Model Weights
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token

model = SourceCodeAuthorCheck().to(device)
model_path = hf_hub_download(repo_id="assix-research/SourceCodeAuthorCheck-SLM-10M", filename="source_code_classifier.pth")
model.load_state_dict(torch.load(model_path, map_location=device, weights_only=True))
model.eval()

# 3. Analyze Code Snippet
code_snippet = "print('Hello World')"
inputs = tokenizer(
    code_snippet, return_tensors="pt", truncation=True, padding="max_length", max_length=1024
).to(device)

with torch.no_grad():
    logits = model(inputs['input_ids'], inputs['attention_mask'])
    prob = torch.sigmoid(logits).item()

print(f"AI Probability: {prob:.1%}")
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using assix-research/SourceCodeAuthorCheck-SLM-10M 1