Expand model card with setup, inference, evaluation and training instructions
Browse files
README.md
CHANGED
|
@@ -13,48 +13,81 @@ tags:
|
|
| 13 |
|
| 14 |
# CADReasoner-CM (cross-modality variant)
|
| 15 |
|
| 16 |
-
CADReasoner-CM is
|
| 17 |
-
|
| 18 |
-
and refines it over
|
| 19 |
-
and the current prediction.
|
| 20 |
|
| 21 |
-
This checkpoint conditions on **both the point cloud and multi-view renders** of the target shape. Combining the
|
| 22 |
-
two modalities improves geometric alignment and recovers fine details that either modality alone tends to miss.
|
| 23 |
|
| 24 |
**Accepted to CVPR 2026 Findings Track**
|
| 25 |
|
| 26 |
**Paper:** https://arxiv.org/abs/2603.29847
|
| 27 |
-
**Code:** https://github.com/zhemdi/CADReasoner
|
| 28 |
**HF paper page:** https://huggingface.co/papers/2603.29847
|
| 29 |
|
| 30 |
## Variants
|
| 31 |
|
| 32 |
-
| Model | Input modality |
|
| 33 |
-
|---|---|
|
| 34 |
-
| [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) | multi-view renders |
|
| 35 |
-
| [`kulibinai/cadreasoner-pc`](https://huggingface.co/kulibinai/cadreasoner-pc) | point cloud |
|
| 36 |
-
| [`kulibinai/cadreasoner-cm`](https://huggingface.co/kulibinai/cadreasoner-cm) | point cloud + multi-view renders
|
| 37 |
|
| 38 |
-
##
|
| 39 |
|
| 40 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
|
| 42 |
```bash
|
| 43 |
-
python3 test.py \
|
| 44 |
--dataset <test_dataset> \
|
| 45 |
--checkpoint kulibinai/cadreasoner-cm \
|
| 46 |
--use_pc true --use_img true \
|
| 47 |
--n_points 128 \
|
| 48 |
-
--n_iters
|
| 49 |
-
--
|
|
|
|
| 50 |
```
|
| 51 |
|
| 52 |
-
|
| 53 |
-
encoder), so it must be loaded through `cadrille.py` from the repository rather than through a
|
| 54 |
-
bare `Qwen2VLForConditionalGeneration.from_pretrained`. `--n_points` must match the value used
|
| 55 |
-
in training (**128**).
|
| 56 |
-
|
| 57 |
-
`<test_dataset>` can be one of:
|
| 58 |
|
| 59 |
* `maksimko123/deepcad_test_mesh`
|
| 60 |
* `maksimko123/fusion360_test_mesh`
|
|
@@ -63,6 +96,80 @@ in training (**128**).
|
|
| 63 |
* `kulibinai/fusion360_test_scan`
|
| 64 |
* `kulibinai/mcb_test_scan`
|
| 65 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
## Citation
|
| 67 |
|
| 68 |
```bibtex
|
|
|
|
| 13 |
|
| 14 |
# CADReasoner-CM (cross-modality variant)
|
| 15 |
|
| 16 |
+
CADReasoner-CM is a variant of [CADReasoner](https://huggingface.co/kulibinai/cadreasoner), a
|
| 17 |
+
vision–language model for **iterative CAD reverse engineering**. It generates a **runnable CadQuery
|
| 18 |
+
program** and refines it over several iterations using geometric feedback from the discrepancy
|
| 19 |
+
between the target shape and the current prediction.
|
| 20 |
|
| 21 |
+
This checkpoint conditions on **both the point cloud and multi-view renders** of the target shape. Combining the two modalities improves geometric alignment and recovers fine details that either modality alone tends to miss.
|
|
|
|
| 22 |
|
| 23 |
**Accepted to CVPR 2026 Findings Track**
|
| 24 |
|
| 25 |
**Paper:** https://arxiv.org/abs/2603.29847
|
| 26 |
+
**Code:** https://github.com/zhemdi/CADReasoner (see the `pc_cm/` directory)
|
| 27 |
**HF paper page:** https://huggingface.co/papers/2603.29847
|
| 28 |
|
| 29 |
## Variants
|
| 30 |
|
| 31 |
+
| Model | Input modality | Flags |
|
| 32 |
+
|---|---|---|
|
| 33 |
+
| [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) | multi-view renders | — (root scripts) |
|
| 34 |
+
| [`kulibinai/cadreasoner-pc`](https://huggingface.co/kulibinai/cadreasoner-pc) | point cloud | `--use_pc true --use_img false` |
|
| 35 |
+
| [`kulibinai/cadreasoner-cm`](https://huggingface.co/kulibinai/cadreasoner-cm) | point cloud + multi-view renders | `--use_pc true --use_img true` |
|
| 36 |
|
| 37 |
+
## Architecture
|
| 38 |
|
| 39 |
+
This checkpoint uses the `Cadrille` architecture: a Qwen2-VL-2B backbone plus a Fourier point-cloud
|
| 40 |
+
encoder. The point features are projected to the hidden size and **scattered into a run of pad
|
| 41 |
+
tokens prepended to the prompt**, so the model cannot be loaded with a bare
|
| 42 |
+
`Qwen2VLForConditionalGeneration.from_pretrained` — `config.json` declares
|
| 43 |
+
`architectures: ["Cadrille"]`, and the class lives in `pc_cm/cadrille.py` in the repository.
|
| 44 |
+
|
| 45 |
+
The point cloud is built from the **discrepancy** between the target mesh and the current
|
| 46 |
+
prediction, not from the target alone: both surfaces are sampled, points whose distance to the
|
| 47 |
+
opposite shape exceeds a percentile threshold are kept, and farthest point sampling reduces them.
|
| 48 |
+
Each direction contributes `n_points` points — `n_points` from GT→pred and `n_points` from pred→GT
|
| 49 |
+
— so the prompt reserves `2 * n_points` pad tokens and each point carries 6 features (position +
|
| 50 |
+
displacement vector). On the first iteration, where no prediction exists yet, the bounding-box
|
| 51 |
+
centre is used in place of the predicted mesh.
|
| 52 |
+
|
| 53 |
+
`--n_points` must match the value used in training: **128**.
|
| 54 |
+
|
| 55 |
+
## Setup
|
| 56 |
+
|
| 57 |
+
```bash
|
| 58 |
+
git clone https://github.com/zhemdi/CADReasoner.git
|
| 59 |
+
cd CADReasoner
|
| 60 |
+
```
|
| 61 |
+
|
| 62 |
+
Build the environment from the provided `Dockerfile`, then add the two packages the `pc_cm/`
|
| 63 |
+
scripts need on top of it:
|
| 64 |
+
|
| 65 |
+
```bash
|
| 66 |
+
pip install opencv-python rtree
|
| 67 |
+
```
|
| 68 |
+
|
| 69 |
+
`opencv-python` is used by the image augmentations, and `rtree` backs trimesh's closest-point
|
| 70 |
+
queries — without it the code silently falls back to a KD-tree over mesh vertices, which is less
|
| 71 |
+
accurate on coarse meshes.
|
| 72 |
+
|
| 73 |
+
The scripts load the model with `attn_implementation="flash_attention_2"` and `torch.bfloat16`, so
|
| 74 |
+
an Ampere-or-newer GPU with `flash-attn` installed is required. Inference shards samples across all
|
| 75 |
+
visible GPUs, one process per GPU, and needs at least one.
|
| 76 |
+
|
| 77 |
+
## Inference
|
| 78 |
|
| 79 |
```bash
|
| 80 |
+
python3 pc_cm/test.py \
|
| 81 |
--dataset <test_dataset> \
|
| 82 |
--checkpoint kulibinai/cadreasoner-cm \
|
| 83 |
--use_pc true --use_img true \
|
| 84 |
--n_points 128 \
|
| 85 |
+
--n_iters 3 \
|
| 86 |
+
--n_samples 4 \
|
| 87 |
+
--outdir preds_cm
|
| 88 |
```
|
| 89 |
|
| 90 |
+
`<test_dataset>` is either a local directory of `.stl` files or a Hugging Face dataset repo id:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
|
| 92 |
* `maksimko123/deepcad_test_mesh`
|
| 93 |
* `maksimko123/fusion360_test_mesh`
|
|
|
|
| 96 |
* `kulibinai/fusion360_test_scan`
|
| 97 |
* `kulibinai/mcb_test_scan`
|
| 98 |
|
| 99 |
+
Each iteration writes a CadQuery program and its compiled mesh per candidate:
|
| 100 |
+
|
| 101 |
+
```text
|
| 102 |
+
preds_cm/<dataset_name>/<iter>/<shape_stem>/<chamfer_distance>.py
|
| 103 |
+
preds_cm/<dataset_name>/<iter>/<shape_stem>/<chamfer_distance>.stl
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
+
Files are renamed to their chamfer distance so the next iteration can pick the best candidates to
|
| 107 |
+
refine. `--n_samples` controls how many candidates are sampled per shape (the first is greedy, the
|
| 108 |
+
rest use temperature 1.2).
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
```bash
|
| 113 |
+
python3 pc_cm/evaluate.py \
|
| 114 |
+
--dataset <test_dataset> \
|
| 115 |
+
--pred_dir preds_cm/<dataset_name>
|
| 116 |
+
```
|
| 117 |
+
|
| 118 |
+
For each refinement iteration this reports median chamfer distance (scaled by 1000), mean IoU, and
|
| 119 |
+
the invalidity ratio — the fraction of shapes with no valid prediction. Each shape is scored with
|
| 120 |
+
the best candidate seen up to that iteration, so the numbers are cumulative across iterations.
|
| 121 |
+
|
| 122 |
+
## Training
|
| 123 |
+
|
| 124 |
+
The curriculum runs over dataset groups `0, 1, 2` (see `data/README.md` for preparing the split).
|
| 125 |
+
Group 0 starts from `Qwen/Qwen2-VL-2B-Instruct`; each later group starts from the previous group's
|
| 126 |
+
final checkpoint.
|
| 127 |
+
|
| 128 |
+
```bash
|
| 129 |
+
torchrun --nproc_per_node <n_gpus> pc_cm/train.py \
|
| 130 |
+
--dataset_dir <train_dataset_dir> \
|
| 131 |
+
--use_pc true --use_img true \
|
| 132 |
+
--n_points 128
|
| 133 |
+
```
|
| 134 |
+
|
| 135 |
+
Checkpoints and logs are written to `runs/<experiment>/<experiment>_<timestamp>/`, with per-group
|
| 136 |
+
weights under `model/<group>/final`. Pass `--run_dir` and `--groups` to resume an interrupted run,
|
| 137 |
+
and `--skip_generate_code` / `--skip_generate_meshes` to reuse refinement samples already on disk.
|
| 138 |
+
|
| 139 |
+
## Loading the model directly
|
| 140 |
+
|
| 141 |
+
If you are integrating the model rather than using the scripts:
|
| 142 |
+
|
| 143 |
+
```python
|
| 144 |
+
import torch
|
| 145 |
+
from transformers import AutoProcessor
|
| 146 |
+
from cadrille import Cadrille # from pc_cm/
|
| 147 |
+
|
| 148 |
+
processor = AutoProcessor.from_pretrained(
|
| 149 |
+
"kulibinai/cadreasoner-cm",
|
| 150 |
+
resized_width=14 * 17 * 2,
|
| 151 |
+
resized_height=14 * 17 * 4,
|
| 152 |
+
padding_side="left",
|
| 153 |
+
use_fast=True,
|
| 154 |
+
)
|
| 155 |
+
model = Cadrille.from_pretrained(
|
| 156 |
+
"kulibinai/cadreasoner-cm",
|
| 157 |
+
torch_dtype=torch.bfloat16,
|
| 158 |
+
attn_implementation="flash_attention_2",
|
| 159 |
+
).cuda().eval()
|
| 160 |
+
```
|
| 161 |
+
|
| 162 |
+
`forward` and `generate` additionally require `point_clouds` (a `(batch, 2 * n_points, 6)` tensor),
|
| 163 |
+
`is_pc` and `is_img` — they are not optional, and the pad-token prefix must already be present in
|
| 164 |
+
the prompt. See `generate_predictions_process` in `pc_cm/test.py` for a complete, working example.
|
| 165 |
+
|
| 166 |
+
## Limitations
|
| 167 |
+
|
| 168 |
+
The image-conditioned [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) gives
|
| 169 |
+
the best results in our implementation; the geometry-conditioned variants are released for
|
| 170 |
+
reproducibility and for settings where rendered views are unavailable. The point-cloud modality also
|
| 171 |
+
proved unstable under RL fine-tuning, so it is not being developed further at present.
|
| 172 |
+
|
| 173 |
## Citation
|
| 174 |
|
| 175 |
```bibtex
|