kulibinai commited on
Commit
f9af3b2
·
verified ·
1 Parent(s): baa670d

Expand model card with setup, inference, evaluation and training instructions

Browse files
Files changed (1) hide show
  1. README.md +130 -23
README.md CHANGED
@@ -13,48 +13,81 @@ tags:
13
 
14
  # CADReasoner-CM (cross-modality variant)
15
 
16
- CADReasoner-CM is the **cross-modality** variant of [CADReasoner](https://huggingface.co/kulibinai/cadreasoner),
17
- a vision–language model for **iterative CAD reverse engineering**. It generates a **runnable CadQuery program**
18
- and refines it over multiple iterations using geometric feedback from the discrepancy between the target shape
19
- and the current prediction.
20
 
21
- This checkpoint conditions on **both the point cloud and multi-view renders** of the target shape. Combining the
22
- two modalities improves geometric alignment and recovers fine details that either modality alone tends to miss.
23
 
24
  **Accepted to CVPR 2026 Findings Track**
25
 
26
  **Paper:** https://arxiv.org/abs/2603.29847
27
- **Code:** https://github.com/zhemdi/CADReasoner
28
  **HF paper page:** https://huggingface.co/papers/2603.29847
29
 
30
  ## Variants
31
 
32
- | Model | Input modality |
33
- |---|---|
34
- | [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) | multi-view renders |
35
- | [`kulibinai/cadreasoner-pc`](https://huggingface.co/kulibinai/cadreasoner-pc) | point cloud |
36
- | [`kulibinai/cadreasoner-cm`](https://huggingface.co/kulibinai/cadreasoner-cm) | point cloud + multi-view renders (cross-modality) |
37
 
38
- ## Usage
39
 
40
- Use the inference and evaluation scripts from the repository:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41
 
42
  ```bash
43
- python3 test.py \
44
  --dataset <test_dataset> \
45
  --checkpoint kulibinai/cadreasoner-cm \
46
  --use_pc true --use_img true \
47
  --n_points 128 \
48
- --n_iters <n_iters> \
49
- --outdir <outdir>
 
50
  ```
51
 
52
- The checkpoint uses the custom `Cadrille` architecture (Qwen2-VL plus a Fourier point-cloud
53
- encoder), so it must be loaded through `cadrille.py` from the repository rather than through a
54
- bare `Qwen2VLForConditionalGeneration.from_pretrained`. `--n_points` must match the value used
55
- in training (**128**).
56
-
57
- `<test_dataset>` can be one of:
58
 
59
  * `maksimko123/deepcad_test_mesh`
60
  * `maksimko123/fusion360_test_mesh`
@@ -63,6 +96,80 @@ in training (**128**).
63
  * `kulibinai/fusion360_test_scan`
64
  * `kulibinai/mcb_test_scan`
65
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
66
  ## Citation
67
 
68
  ```bibtex
 
13
 
14
  # CADReasoner-CM (cross-modality variant)
15
 
16
+ CADReasoner-CM is a variant of [CADReasoner](https://huggingface.co/kulibinai/cadreasoner), a
17
+ vision–language model for **iterative CAD reverse engineering**. It generates a **runnable CadQuery
18
+ program** and refines it over several iterations using geometric feedback from the discrepancy
19
+ between the target shape and the current prediction.
20
 
21
+ This checkpoint conditions on **both the point cloud and multi-view renders** of the target shape. Combining the two modalities improves geometric alignment and recovers fine details that either modality alone tends to miss.
 
22
 
23
  **Accepted to CVPR 2026 Findings Track**
24
 
25
  **Paper:** https://arxiv.org/abs/2603.29847
26
+ **Code:** https://github.com/zhemdi/CADReasoner (see the `pc_cm/` directory)
27
  **HF paper page:** https://huggingface.co/papers/2603.29847
28
 
29
  ## Variants
30
 
31
+ | Model | Input modality | Flags |
32
+ |---|---|---|
33
+ | [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) | multi-view renders | — (root scripts) |
34
+ | [`kulibinai/cadreasoner-pc`](https://huggingface.co/kulibinai/cadreasoner-pc) | point cloud | `--use_pc true --use_img false` |
35
+ | [`kulibinai/cadreasoner-cm`](https://huggingface.co/kulibinai/cadreasoner-cm) | point cloud + multi-view renders | `--use_pc true --use_img true` |
36
 
37
+ ## Architecture
38
 
39
+ This checkpoint uses the `Cadrille` architecture: a Qwen2-VL-2B backbone plus a Fourier point-cloud
40
+ encoder. The point features are projected to the hidden size and **scattered into a run of pad
41
+ tokens prepended to the prompt**, so the model cannot be loaded with a bare
42
+ `Qwen2VLForConditionalGeneration.from_pretrained` — `config.json` declares
43
+ `architectures: ["Cadrille"]`, and the class lives in `pc_cm/cadrille.py` in the repository.
44
+
45
+ The point cloud is built from the **discrepancy** between the target mesh and the current
46
+ prediction, not from the target alone: both surfaces are sampled, points whose distance to the
47
+ opposite shape exceeds a percentile threshold are kept, and farthest point sampling reduces them.
48
+ Each direction contributes `n_points` points — `n_points` from GT→pred and `n_points` from pred→GT
49
+ — so the prompt reserves `2 * n_points` pad tokens and each point carries 6 features (position +
50
+ displacement vector). On the first iteration, where no prediction exists yet, the bounding-box
51
+ centre is used in place of the predicted mesh.
52
+
53
+ `--n_points` must match the value used in training: **128**.
54
+
55
+ ## Setup
56
+
57
+ ```bash
58
+ git clone https://github.com/zhemdi/CADReasoner.git
59
+ cd CADReasoner
60
+ ```
61
+
62
+ Build the environment from the provided `Dockerfile`, then add the two packages the `pc_cm/`
63
+ scripts need on top of it:
64
+
65
+ ```bash
66
+ pip install opencv-python rtree
67
+ ```
68
+
69
+ `opencv-python` is used by the image augmentations, and `rtree` backs trimesh's closest-point
70
+ queries — without it the code silently falls back to a KD-tree over mesh vertices, which is less
71
+ accurate on coarse meshes.
72
+
73
+ The scripts load the model with `attn_implementation="flash_attention_2"` and `torch.bfloat16`, so
74
+ an Ampere-or-newer GPU with `flash-attn` installed is required. Inference shards samples across all
75
+ visible GPUs, one process per GPU, and needs at least one.
76
+
77
+ ## Inference
78
 
79
  ```bash
80
+ python3 pc_cm/test.py \
81
  --dataset <test_dataset> \
82
  --checkpoint kulibinai/cadreasoner-cm \
83
  --use_pc true --use_img true \
84
  --n_points 128 \
85
+ --n_iters 3 \
86
+ --n_samples 4 \
87
+ --outdir preds_cm
88
  ```
89
 
90
+ `<test_dataset>` is either a local directory of `.stl` files or a Hugging Face dataset repo id:
 
 
 
 
 
91
 
92
  * `maksimko123/deepcad_test_mesh`
93
  * `maksimko123/fusion360_test_mesh`
 
96
  * `kulibinai/fusion360_test_scan`
97
  * `kulibinai/mcb_test_scan`
98
 
99
+ Each iteration writes a CadQuery program and its compiled mesh per candidate:
100
+
101
+ ```text
102
+ preds_cm/<dataset_name>/<iter>/<shape_stem>/<chamfer_distance>.py
103
+ preds_cm/<dataset_name>/<iter>/<shape_stem>/<chamfer_distance>.stl
104
+ ```
105
+
106
+ Files are renamed to their chamfer distance so the next iteration can pick the best candidates to
107
+ refine. `--n_samples` controls how many candidates are sampled per shape (the first is greedy, the
108
+ rest use temperature 1.2).
109
+
110
+ ## Evaluation
111
+
112
+ ```bash
113
+ python3 pc_cm/evaluate.py \
114
+ --dataset <test_dataset> \
115
+ --pred_dir preds_cm/<dataset_name>
116
+ ```
117
+
118
+ For each refinement iteration this reports median chamfer distance (scaled by 1000), mean IoU, and
119
+ the invalidity ratio — the fraction of shapes with no valid prediction. Each shape is scored with
120
+ the best candidate seen up to that iteration, so the numbers are cumulative across iterations.
121
+
122
+ ## Training
123
+
124
+ The curriculum runs over dataset groups `0, 1, 2` (see `data/README.md` for preparing the split).
125
+ Group 0 starts from `Qwen/Qwen2-VL-2B-Instruct`; each later group starts from the previous group's
126
+ final checkpoint.
127
+
128
+ ```bash
129
+ torchrun --nproc_per_node <n_gpus> pc_cm/train.py \
130
+ --dataset_dir <train_dataset_dir> \
131
+ --use_pc true --use_img true \
132
+ --n_points 128
133
+ ```
134
+
135
+ Checkpoints and logs are written to `runs/<experiment>/<experiment>_<timestamp>/`, with per-group
136
+ weights under `model/<group>/final`. Pass `--run_dir` and `--groups` to resume an interrupted run,
137
+ and `--skip_generate_code` / `--skip_generate_meshes` to reuse refinement samples already on disk.
138
+
139
+ ## Loading the model directly
140
+
141
+ If you are integrating the model rather than using the scripts:
142
+
143
+ ```python
144
+ import torch
145
+ from transformers import AutoProcessor
146
+ from cadrille import Cadrille # from pc_cm/
147
+
148
+ processor = AutoProcessor.from_pretrained(
149
+ "kulibinai/cadreasoner-cm",
150
+ resized_width=14 * 17 * 2,
151
+ resized_height=14 * 17 * 4,
152
+ padding_side="left",
153
+ use_fast=True,
154
+ )
155
+ model = Cadrille.from_pretrained(
156
+ "kulibinai/cadreasoner-cm",
157
+ torch_dtype=torch.bfloat16,
158
+ attn_implementation="flash_attention_2",
159
+ ).cuda().eval()
160
+ ```
161
+
162
+ `forward` and `generate` additionally require `point_clouds` (a `(batch, 2 * n_points, 6)` tensor),
163
+ `is_pc` and `is_img` — they are not optional, and the pad-token prefix must already be present in
164
+ the prompt. See `generate_predictions_process` in `pc_cm/test.py` for a complete, working example.
165
+
166
+ ## Limitations
167
+
168
+ The image-conditioned [`kulibinai/cadreasoner`](https://huggingface.co/kulibinai/cadreasoner) gives
169
+ the best results in our implementation; the geometry-conditioned variants are released for
170
+ reproducibility and for settings where rendered views are unavailable. The point-cloud modality also
171
+ proved unstable under RL fine-tuning, so it is not being developed further at present.
172
+
173
  ## Citation
174
 
175
  ```bibtex