Text Generation
Transformers
Safetensors
llada2_moe
dllm
diffusion
llm
text_generation
conversational
custom_code
Instructions to use inclusionAI/LLaDA2.2-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use inclusionAI/LLaDA2.2-flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="inclusionAI/LLaDA2.2-flash", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("inclusionAI/LLaDA2.2-flash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use inclusionAI/LLaDA2.2-flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "inclusionAI/LLaDA2.2-flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inclusionAI/LLaDA2.2-flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/inclusionAI/LLaDA2.2-flash
- SGLang
How to use inclusionAI/LLaDA2.2-flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "inclusionAI/LLaDA2.2-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inclusionAI/LLaDA2.2-flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "inclusionAI/LLaDA2.2-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "inclusionAI/LLaDA2.2-flash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use inclusionAI/LLaDA2.2-flash with Docker Model Runner:
docker model run hf.co/inclusionAI/LLaDA2.2-flash
[fix] gate M2T unmasking on sampled-token confidence p_s
Browse files- modeling_llada2_moe.py +4 -4
modeling_llada2_moe.py
CHANGED
|
@@ -1520,14 +1520,14 @@ class LLaDA2MoeModelLM(LLaDA2MoePreTrainedModel, GenerationMixin):
|
|
| 1520 |
neg_inf = torch.full_like(p0, -float("inf"))
|
| 1521 |
|
| 1522 |
# 3. M2T (mask -> token): threshold-gated with a per-step floor.
|
| 1523 |
-
#
|
| 1524 |
-
# ``
|
| 1525 |
-
#
|
| 1526 |
mt2_index = torch.zeros(block_length, dtype=torch.bool, device=device)
|
| 1527 |
if mask_index.any():
|
| 1528 |
if step_id < len(transfer_schedule):
|
| 1529 |
num_need = transfer_schedule[step_id].item() + new_mask_count
|
| 1530 |
-
mask_conf = torch.where(mask_index,
|
| 1531 |
high_conf = (mask_conf > threshold) & mask_index
|
| 1532 |
if high_conf.sum().item() >= num_need:
|
| 1533 |
mt2_index = high_conf
|
|
|
|
| 1520 |
neg_inf = torch.full_like(p0, -float("inf"))
|
| 1521 |
|
| 1522 |
# 3. M2T (mask -> token): threshold-gated with a per-step floor.
|
| 1523 |
+
# The gate uses ``p_s`` -- the confidence of the token that will actually be written
|
| 1524 |
+
# (``x_s``) -- so a position is only unmasked when the sampled token itself is
|
| 1525 |
+
# confident. When temperature == 0, ``p_s == p0``, so this reduces to greedy behavior.
|
| 1526 |
mt2_index = torch.zeros(block_length, dtype=torch.bool, device=device)
|
| 1527 |
if mask_index.any():
|
| 1528 |
if step_id < len(transfer_schedule):
|
| 1529 |
num_need = transfer_schedule[step_id].item() + new_mask_count
|
| 1530 |
+
mask_conf = torch.where(mask_index, p_s, neg_inf)
|
| 1531 |
high_conf = (mask_conf > threshold) & mask_index
|
| 1532 |
if high_conf.sum().item() >= num_need:
|
| 1533 |
mt2_index = high_conf
|