AI & ML interests

None defined yet.

Recent Activity

Pranav2748Β  updated a bucket about 8 hours ago
Mercity/SkillsStorage
Rishik001Β  published a dataset about 20 hours ago
Mercity/sop-compliance-dpo-pairs
Rishik001Β  updated a dataset 1 day ago
Mercity/sop-compliance-dpo-pairs
View all activity

Organization Card

Mercity

A research-first firm building custom data, models and evaluation
for problems an off-the-shelf API cannot solve.

Website Research GitHub LinkedIn


πŸ‘‹ About us

Mercity Research Labs works with product companies and enterprises on custom training, new model architectures, optimization and private research. We handle every stage, from the training data to the evaluation, and our clients own the weights and the code.

We are research-first and work in the open. The datasets, weights and tools on this page come straight out of our own projects, published so anyone can use them and reproduce our results.

πŸ§ͺ What we focus on

Synthetic data generation. We have generated hundreds of billions of tokens of synthetic data and trained models on it. Every dataset ships with its generation code, filters and a coverage report.

Custom architecture and training. When an LLM or API model isn't the right fit, we design and train models around the task, including modified attention, custom kernels, pruning and distillation (for example, Qwen3-8B pruned from 36 layers to 30).

Evaluation and benchmarking. Every project starts with a domain-specific benchmark, and every change is measured against it. We read the outputs ourselves.

πŸ“¦ Highlights on the Hub

πŸ›‘οΈ Topical Guardrails SAE-based classifiers on Gemma 1B / 4B for keeping models on topic
⛸️ Figure Skating Classification Five pose-based architectures (CNN+BiLSTM, Transformer, CTM, Transformer+BiLSTM, GCN) trained on SkatingVerse
✍️ creative-writing-llm Text-generation model tuned for creative writing
🧠 ReasonBridge-URT ~1M-row reasoning dataset
🌐 FineWeb-Preprocessed 10M-document preprocessed pretraining corpus
πŸ–ΌοΈ laion-subset Curated image-text subset of LAION

Browse everything: models Β· datasets Β· collections

πŸ› οΈ Open-source tools

Tool Stack What it does
Simula Python Deliberate synthetic data generation: describe the distribution as a spec, and Simula plans, generates and verifies a corpus against it
PromptKeep Python Versioned prompt templates with lineage, plus a drop-in OpenAI wrapper that records every run

🀝 Work with us

Have a problem that doesn't fit the API? If an off-the-shelf model gets you most of the way and the rest is what your business depends on, we should talk.

Start a conversation β†’