chanyoungkim commited on
Commit
f6aa369
·
verified ·
1 Parent(s): 84a93ef

Add model, config, tokenizer and model card

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,192 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-1.5B-Instruct
4
+ library_name: transformers
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - reinforcement-learning
8
+ - agentic-rl
9
+ - alfworld
10
+ - strata
11
+ - grpo
12
+ - qwen2.5
13
+ ---
14
+
15
+ # StraTA / Qwen2.5-1.5B-Instruct / ALFWorld — optimizer step 120
16
+
17
+ Reproduction of **StraTA** (Strategic Trajectory Abstraction,
18
+ [arXiv:2605.06642](https://arxiv.org/abs/2605.06642), code
19
+ [xxyQwQ/StraTA](https://github.com/xxyQwQ/StraTA) @ `592238b`) on ALFWorld with
20
+ `Qwen/Qwen2.5-1.5B-Instruct`.
21
+
22
+ **This checkpoint: optimizer step 120 of 150.**
23
+
24
+ | | |
25
+ |---|---|
26
+ | Optimizer step | **120** / 150 |
27
+ | Validation success rate at this step | **69.29%** (ALFWorld `valid_seen`, 140 tasks) |
28
+ | Base model success rate | 0.71% |
29
+ | Backbone | `Qwen/Qwen2.5-1.5B-Instruct` (1.54B params, bf16) |
30
+ | Environment | ALFWorld (AgentGym HTTP interface) |
31
+ | Hardware | 2x NVIDIA A100 80GB |
32
+ | Wall-clock (full 150 steps) | ~44 hours |
33
+
34
+ ## Training curve
35
+
36
+ Validation success rate on the full ALFWorld `valid_seen` split (140 tasks),
37
+ measured every 10 optimizer steps during training. Step 0 is the untrained base model.
38
+
39
+ | optimizer step | success rate (%) |
40
+ |---|---|
41
+ | 0 | 0.71 |
42
+ | 10 | 0.71 |
43
+ | 20 | 3.57 |
44
+ | 30 | 1.43 |
45
+ | 40 | 14.29 |
46
+ | 50 | 18.57 |
47
+ | 60 | 6.43 |
48
+ | 70 | 22.86 |
49
+ | 80 | 37.14 |
50
+ | 90 | 42.14 |
51
+ | 100 | 43.57 |
52
+ | 110 | 50.71 |
53
+ | 120 | 69.29 |
54
+ | 130 | 85.00 |
55
+ | 140 | 83.57 |
56
+ | 150 | 86.43 |
57
+
58
+ ## Final evaluation (step 150, 2 seeds)
59
+
60
+ Run with `evaluate-strata-alfworld-text-qwen2.5-1.5b.sh`: `val_only`, 140 tasks,
61
+ temperature 0.7 / top_p 0.8 / top_k 20, max 50 env steps.
62
+
63
+ | subtask | seed 1 | seed 2 | mean ± std |
64
+ |---|---|---|---|
65
+ | pick | 97.14 | 97.14 | **97.14 ± 0.00** |
66
+ | clean | 92.59 | 92.59 | **92.59 ± 0.00** |
67
+ | pick2 | 91.67 | 87.50 | **89.58 ± 2.95** |
68
+ | heat | 81.25 | 81.25 | **81.25 ± 0.00** |
69
+ | cool | 76.00 | 72.00 | **74.00 ± 2.83** |
70
+ | look | 76.92 | 69.23 | **73.08 ± 5.44** |
71
+ | ALL | 87.86 | 85.71 | **86.79 ± 1.52** |
72
+
73
+ Auxiliary metrics: action `format_rate` **1.000** (both seeds), mean episode length
74
+ **13.2–14.2** steps (base model: 49.7).
75
+
76
+ ## Hyperparameters
77
+
78
+ All algorithm hyperparameters follow the paper: 150 steps, batch size 16, 4 strategies
79
+ per task, 8 rollouts per strategy, oversampling ratio sigma=8, aggregation ratio
80
+ delta=0.5, length penalty threshold lambda=0.5, self-judgment reward weight kappa=0.1.
81
+
82
+ | group | parameter | value |
83
+ |---|---|---|
84
+ | StraTA | `num_stras_per_run` (strategies per task) | 4 |
85
+ | StraTA | `num_trajs_per_stra` (rollouts per strategy) | 8 |
86
+ | StraTA | `num_cands_per_stra` (oversampling ratio, sigma) | 8 |
87
+ | StraTA | `delta_coef` (aggregation ratio) | 0.5 |
88
+ | StraTA | `lambda_coef` (length penalty threshold) | 0.5 |
89
+ | StraTA | `kappa_coef` (self-judgment reward weight) | 0.1 |
90
+ | StraTA | `use_diverse_stras` / `use_self_judgment` | True / True |
91
+ | StraTA | `use_length_reward` / `use_format_reward` | True / True |
92
+ | StraTA | strategy embedding model | `sentence-transformers/all-MiniLM-L6-v2` (CPU) |
93
+ | RL | `adv_estimator` | grpo |
94
+ | RL | `algorithm.gamma` | 0 |
95
+ | RL | `kl_ctrl.kl_coef` | 0.001 |
96
+ | RL | `actor.use_kl_loss` | False |
97
+ | RL | `clip_ratio_high` | 0.28 |
98
+ | RL | `loss_agg_mode` | seq-mean-token-sum |
99
+ | RL | `entropy_coeff` | 0 |
100
+ | RL | `stepwise_advantage` | enabled, `per_step` |
101
+ | Optim | learning rate | 1e-6 |
102
+ | Optim | `ppo_mini_batch_size` (global) | 1024 |
103
+ | Optim | `use_dynamic_bsz` / `ppo_max_token_len_per_gpu` | True / 32768 |
104
+ | Data | `train_batch_size` (tasks per step) | 16 |
105
+ | Data | episodes per step | 16 x 4 x 8 = **512** |
106
+ | Data | `max_prompt_length` / `max_response_length` | 7168 / 1024 |
107
+ | Data | env `num_histories` | 2 |
108
+ | Data | max env steps per episode | 50 |
109
+ | Rollout | engine / temperature | vLLM 0.11.0 (async) / 1.0 |
110
+ | Eval | `val_kwargs` | n=1, temperature 0.7, top_p 0.8, top_k 20 |
111
+ | Train | total steps | 150 (`total_training_steps=151`) |
112
+
113
+
114
+ ### Infrastructure deltas
115
+
116
+ This run reproduces the upstream 7B recipe on 2x A100 80GB. A full diff of the
117
+ resolved Hydra config against `train-strata-alfworld-text-qwen2.5-7b.sh` shows
118
+ **6 differing keys out of 533**, all infrastructure -- no algorithm or optimization
119
+ parameter was changed:
120
+
121
+ | key | upstream 7B | this run | effect on results |
122
+ |---|---|---|---|
123
+ | `trainer.n_gpus_per_node` | 8 | 2 | none -- verl computes `ppo_mini_batch_size *= rollout.n` then `//= world_size`, so the global mini-batch is invariant |
124
+ | `actor.fsdp_config.param_offload` | True | False | none -- CPU offload only |
125
+ | `actor.fsdp_config.optimizer_offload` | True | False | none |
126
+ | `ref.fsdp_config.param_offload` | True | False | none |
127
+ | `actor.checkpoint.save_contents` | `[hf_model]` | `[model,optimizer,extra,hf_model]` | none -- what gets written |
128
+ | `trainer.save_freq` | 50 | 10 | none -- checkpoint cadence |
129
+
130
+
131
+ ## Usage
132
+
133
+ ```python
134
+ import torch
135
+ from transformers import AutoModelForCausalLM, AutoTokenizer
136
+
137
+ repo = "chanyoungkim/strata-qwen2.5-1.5b-alfworld-step120"
138
+ model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
139
+ tok = AutoTokenizer.from_pretrained(repo)
140
+ ```
141
+
142
+ The model expects the StraTA ALFWorld prompt format: it first emits a plan inside
143
+ `<strategy>...</strategy>`, then acts, emitting one action per turn inside
144
+ `<action>...</action>`.
145
+
146
+ ## Files
147
+
148
+ | path | contents |
149
+ |---|---|
150
+ | `model.safetensors`, `config.json`, tokenizer files | bf16 weights, ready for `from_pretrained`. `lm_head` is omitted because `tie_word_embeddings=true` (same layout as the base Qwen release). |
151
+ | `training_state/model_world_size_2_rank_{0,1}.pt` | FSDP-sharded fp32 weights |
152
+ | `training_state/optim_world_size_2_rank_{0,1}.pt` | Adam optimizer state |
153
+ | `training_state/extra_state_world_size_2_rank_{0,1}.pt` | RNG and scheduler state |
154
+ | `training_state/fsdp_config.json` | FSDP layout descriptor |
155
+
156
+ `training_state/` lets you resume RL training from this exact optimizer step. Note the
157
+ shards are written for **`world_size=2`** (2 GPUs, FSDP). Resuming on a different number
158
+ of GPUs requires resharding. To resume, place the files back as
159
+ `checkpoints/<exp>/models/global_step_120/actor/` and point verl at that directory.
160
+
161
+ The published `model.safetensors` is a bf16 cast of the fp32 master weights; the
162
+ original fp32 tensors are preserved inside `training_state/`.
163
+
164
+ ## All checkpoints from this run
165
+
166
+ [step 10](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step010) · [step 20](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step020) · [step 30](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step030) · [step 40](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step040) · [step 50](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step050) · [step 60](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step060) · [step 70](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step070) · [step 80](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step080) · [step 90](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step090) · [step 100](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step100) · [step 110](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step110) · [step 120](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step120) · [step 130](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step130) · [step 140](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step140) · [step 150](https://huggingface.co/chanyoungkim/strata-qwen2.5-1.5b-alfworld-step150)
167
+
168
+ ## Caveats
169
+
170
+ - **Do not use `is_correct` or `pass@1`** from the logs. Both are hardcoded to `1.0`
171
+ in the workflow and are meaningless. The correct metric is
172
+ `val/average/metrics/success_rate`.
173
+ - Evaluation uses the StraTA inference protocol (one strategy, then one trajectory
174
+ conditioned on it; no diversity selection, no self-judgment). This is what
175
+ `strata_workflow.run()` does when `rollout_engine.validate` is set, and it matches
176
+ `evaluate-strata-*.sh`.
177
+ - Numbers here are **not** directly comparable to models evaluated under a different
178
+ agent scaffold (different prompt format, context window, or sampling temperature).
179
+ - The evaluation set is ALFWorld `valid_seen` (140 tasks, all of them -- not a
180
+ subsample). `valid_unseen` was not evaluated.
181
+
182
+
183
+ ## Citation
184
+
185
+ ```bibtex
186
+ @article{xue2026strata,
187
+ title={StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction},
188
+ author={Xue, Xiangyuan and Zhou, Yifan and Wang, Zidong and Tang, Shengji and Torr, Philip and Ouyang, Wanli and Bai, Lei and Yin, Zhenfei},
189
+ journal={arXiv preprint arXiv:2605.06642},
190
+ year={2026}
191
+ }
192
+ ```
added_tokens.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "</tool_call>": 151658,
3
+ "<tool_call>": 151657,
4
+ "<|box_end|>": 151649,
5
+ "<|box_start|>": 151648,
6
+ "<|endoftext|>": 151643,
7
+ "<|file_sep|>": 151664,
8
+ "<|fim_middle|>": 151660,
9
+ "<|fim_pad|>": 151662,
10
+ "<|fim_prefix|>": 151659,
11
+ "<|fim_suffix|>": 151661,
12
+ "<|im_end|>": 151645,
13
+ "<|im_start|>": 151644,
14
+ "<|image_pad|>": 151655,
15
+ "<|object_ref_end|>": 151647,
16
+ "<|object_ref_start|>": 151646,
17
+ "<|quad_end|>": 151651,
18
+ "<|quad_start|>": 151650,
19
+ "<|repo_name|>": 151663,
20
+ "<|video_pad|>": 151656,
21
+ "<|vision_end|>": 151653,
22
+ "<|vision_pad|>": 151654,
23
+ "<|vision_start|>": 151652
24
+ }
chat_template.jinja ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0]['role'] == 'system' %}
4
+ {{- messages[0]['content'] }}
5
+ {%- else %}
6
+ {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
7
+ {%- endif %}
8
+ {{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
9
+ {%- for tool in tools %}
10
+ {{- "\n" }}
11
+ {{- tool | tojson }}
12
+ {%- endfor %}
13
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
14
+ {%- else %}
15
+ {%- if messages[0]['role'] == 'system' %}
16
+ {{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
17
+ {%- else %}
18
+ {{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
19
+ {%- endif %}
20
+ {%- endif %}
21
+ {%- for message in messages %}
22
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
23
+ {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
24
+ {%- elif message.role == "assistant" %}
25
+ {{- '<|im_start|>' + message.role }}
26
+ {%- if message.content %}
27
+ {{- '\n' + message.content }}
28
+ {%- endif %}
29
+ {%- for tool_call in message.tool_calls %}
30
+ {%- if tool_call.function is defined %}
31
+ {%- set tool_call = tool_call.function %}
32
+ {%- endif %}
33
+ {{- '\n<tool_call>\n{"name": "' }}
34
+ {{- tool_call.name }}
35
+ {{- '", "arguments": ' }}
36
+ {{- tool_call.arguments | tojson }}
37
+ {{- '}\n</tool_call>' }}
38
+ {%- endfor %}
39
+ {{- '<|im_end|>\n' }}
40
+ {%- elif message.role == "tool" %}
41
+ {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
42
+ {{- '<|im_start|>user' }}
43
+ {%- endif %}
44
+ {{- '\n<tool_response>\n' }}
45
+ {{- message.content }}
46
+ {{- '\n</tool_response>' }}
47
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
48
+ {{- '<|im_end|>\n' }}
49
+ {%- endif %}
50
+ {%- endif %}
51
+ {%- endfor %}
52
+ {%- if add_generation_prompt %}
53
+ {{- '<|im_start|>assistant\n' }}
54
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen2ForCausalLM"
4
+ ],
5
+ "attention_dropout": 0.0,
6
+ "dtype": "bfloat16",
7
+ "eos_token_id": 151645,
8
+ "hidden_act": "silu",
9
+ "hidden_size": 1536,
10
+ "initializer_range": 0.02,
11
+ "intermediate_size": 8960,
12
+ "layer_types": [
13
+ "full_attention",
14
+ "full_attention",
15
+ "full_attention",
16
+ "full_attention",
17
+ "full_attention",
18
+ "full_attention",
19
+ "full_attention",
20
+ "full_attention",
21
+ "full_attention",
22
+ "full_attention",
23
+ "full_attention",
24
+ "full_attention",
25
+ "full_attention",
26
+ "full_attention",
27
+ "full_attention",
28
+ "full_attention",
29
+ "full_attention",
30
+ "full_attention",
31
+ "full_attention",
32
+ "full_attention",
33
+ "full_attention",
34
+ "full_attention",
35
+ "full_attention",
36
+ "full_attention",
37
+ "full_attention",
38
+ "full_attention",
39
+ "full_attention",
40
+ "full_attention"
41
+ ],
42
+ "max_position_embeddings": 32768,
43
+ "max_window_layers": 21,
44
+ "model_type": "qwen2",
45
+ "num_attention_heads": 12,
46
+ "num_hidden_layers": 28,
47
+ "num_key_value_heads": 2,
48
+ "pad_token_id": 151643,
49
+ "rms_norm_eps": 1e-06,
50
+ "rope_scaling": null,
51
+ "rope_theta": 1000000.0,
52
+ "sliding_window": null,
53
+ "tie_word_embeddings": true,
54
+ "transformers_version": "4.57.3",
55
+ "use_cache": true,
56
+ "use_sliding_window": false,
57
+ "vocab_size": 151936,
58
+ "torch_dtype": "bfloat16"
59
+ }
generation_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 151643,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 151645,
6
+ 151643
7
+ ],
8
+ "pad_token_id": 151643,
9
+ "repetition_penalty": 1.1,
10
+ "temperature": 0.7,
11
+ "top_k": 20,
12
+ "top_p": 0.8,
13
+ "transformers_version": "4.57.3"
14
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:89b0f9aaaa67772e7da25349dc2a48db97e1a8c03e2ca1a1aeb88c8e9fd25fe1
3
+ size 3087467144
special_tokens_map.json ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|im_start|>",
4
+ "<|im_end|>",
5
+ "<|object_ref_start|>",
6
+ "<|object_ref_end|>",
7
+ "<|box_start|>",
8
+ "<|box_end|>",
9
+ "<|quad_start|>",
10
+ "<|quad_end|>",
11
+ "<|vision_start|>",
12
+ "<|vision_end|>",
13
+ "<|vision_pad|>",
14
+ "<|image_pad|>",
15
+ "<|video_pad|>"
16
+ ],
17
+ "eos_token": {
18
+ "content": "<|im_end|>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ },
24
+ "pad_token": {
25
+ "content": "<|endoftext|>",
26
+ "lstrip": false,
27
+ "normalized": false,
28
+ "rstrip": false,
29
+ "single_word": false
30
+ }
31
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9c5ae00e602b8860cbd784ba82a8aa14e8feecec692e7076590d014d7b7fdafa
3
+ size 11421896
tokenizer_config.json ADDED
@@ -0,0 +1,207 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ }
181
+ },
182
+ "additional_special_tokens": [
183
+ "<|im_start|>",
184
+ "<|im_end|>",
185
+ "<|object_ref_start|>",
186
+ "<|object_ref_end|>",
187
+ "<|box_start|>",
188
+ "<|box_end|>",
189
+ "<|quad_start|>",
190
+ "<|quad_end|>",
191
+ "<|vision_start|>",
192
+ "<|vision_end|>",
193
+ "<|vision_pad|>",
194
+ "<|image_pad|>",
195
+ "<|video_pad|>"
196
+ ],
197
+ "bos_token": null,
198
+ "clean_up_tokenization_spaces": false,
199
+ "eos_token": "<|im_end|>",
200
+ "errors": "replace",
201
+ "extra_special_tokens": {},
202
+ "model_max_length": 131072,
203
+ "pad_token": "<|endoftext|>",
204
+ "split_special_tokens": false,
205
+ "tokenizer_class": "Qwen2Tokenizer",
206
+ "unk_token": null
207
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff