Julia router
Resident inference for Julia on CPU or CUDA, located in the existing julia
package. No separate repository. Python loads Bend-generated C with ctypes;
there is no compiler or subprocess in the request path.
Build and use
From the Julia-1 directory (the outer workspace also has an older package
named julia, so running there can import the wrong one):
python -m pip install -e .
python -m julia.router.build --bend /home/klee/.bend/bin/bend
The native build requires Bend 2.0.27, Clang, and Linux. It checks native/PROOF.bend, generated
function arities, and the borrowed-tree contract before compiling. --native-cpu
adds host-specific machine instructions; use that build only on compatible CPUs. The bridge
uses private Bend runtime internals, so upgrading Bend requires adapting and
retesting the bridge. --output selects another library location; pass that
path as library= or set JULIA_ROUTER_LIBRARY at runtime.
from julia.router import FastEngine
engine = FastEngine('/path/to/real/checkpoint', device='cpu', batch_size=16)
# device='cuda' uses CUDA/BF16 and device-side softmax/argmax.
rows = [{
'state': 'Preciso trocar minha senha.',
'question': 'Qual é a intenção?',
'options': ['Redefinir senha', 'Cancelar conta', 'Consultar saldo'],
}]
print(engine.predict(rows))
print(engine.predict(rows, probabilities=False)) # return only indices
Use real checkpoint files, not Git LFS pointer files. The 2026-09-23 runtime audit loaded the actual 551 MiB checkpoint offline and exercised resident inference.
Faster inference path
- Bounded LRU caches reuse token fragments and full request encodings. Model forward passes still run on every request; no answer cache inflates timings.
- Sort by encoded length before bounded microbatches, then restore request order. This reduces padding on mixed-length inputs; it does not guarantee a speedup for every batch or request size.
- Build one NumPy arena for integer inputs, with zero-copy CPU tensor views. CUDA uses two pinned bulk transfers instead of per-field transfers.
- Softmax and argmax run on the model's device. CUDA does not copy hidden states or attention matrices into CPU Bend kernels.
compile_model=Trueenables optionaltorch.compile(dynamic=True). Compilation adds first-call cost and requires validation on the target device.- CPU inference retains private file-backed safetensors storage instead of copying
the full vocabulary embedding into anonymous RAM. Keep checkpoint files immutable
while loaded; use
memory_map=Falsefor a detached copy. Training model loading still copies by default. - CPU PyTorch heads compute only option queries and feed-forward outputs in the
final layer. Full context keys/values are retained.
marker_only_head=Falseselects the original path; training and the separate experimental Bend normalization head use the original path automatically. CUDA retains the original default until hardware validation. strict_encoding=Truerejects marker injection and any question/option/state truncation.encoding_info(rows)audits the same cached encoding used in inference; the game worker no longer tokenizes each request twice.- Preserve Julia's original marker serialization, truncation, weights, and
option order. Returned probabilities use the model card presentation rules;
logits()retains raw scores.
CPU defaults to Python/PyTorch (torch), with no Bend build required. CUDA also
uses PyTorch. Select transformer_backend='bend-dense' explicitly to use the
optional CPU FP32 Bend encoder. compile_model=True works with the default Torch
backend.
Bend transformer operations
native/router.bend implements candidate argmax, numerically stabilized softmax,
and two-pass LayerNorm (mean, centered variance, affine scale/bias). Balanced
Leaf/Fork trees expose candidate reductions. LayerNorm forks over independent
rows and uses flat tail loops over feature lists within each row, following the
Bend guide's coarse-work/flat-leaf cost model. All use Bend's own heap. The C adapter transports arrays and owns the ABI; it does
not duplicate the numerical algorithms.
from julia.router import BendReducer, FastEngine
bend = BendReducer()
index, probabilities = bend.softmax([1.0, 3.0, -2.0])
normalized = bend.layernorm([[1.0, 2.0, 3.0]])
# Explicit experimental CPU transformer-head backend:
engine = FastEngine('/path/to/checkpoint', device='cpu',
transformer_backend='bend', bend_postprocess=True)
# All 88 encoder projections of the real checkpoint also execute in Bend:
# Set JULIA_BEND_THREADS=4 before creating the engine.
engine = FastEngine('/path/to/checkpoint', device='cpu',
transformer_backend='bend-dense', bend_postprocess=True)
The experimental head executes its pre-attention and pre-MLP LayerNorms, plus
scorer LayerNorm, in Bend. Dense projections, attention, and the encoder remain
PyTorch. This is not a complete transformer rewrite in Bend. Bend head math
is inference-only CPU float32. The CPU library defaults to up to eight available Bend runtime
workers. Set JULIA_BEND_THREADS=4 before the first native call to test row
parallelism; the worker count is fixed for the process. Pool sizing depends on the kernel and workload; the older normalization-only
experiment does not determine the dense backend defaults. Reduction order differs from the
old tree algorithm and PyTorch: probabilities can differ by floating-point rounding,
and effectively tied choices may choose a different index.
Calls share a mutex because Bend's runtime has global state. Instantiate model
workers in spawned processes; do not fork an active inference process.
Resident dense projections (CPU default)
bend-dense executes all 88 encoder projections in Bend. Resident packed FP32
arrays replace linked weight trees. Contiguous 8×8 tiles expose vector arithmetic;
coarse parallel ranges end in flat tail loops. Shared input/weight handles are
read-only, output tiles are disjoint, and all handles are joined after evaluation.
The bridge packs/transports buffers and checks bounds; the public API validates
finite values, while engine inference validates weights once and final logits.
PyTorch retains embeddings, SDPA, elementwise activations and the selected-output
decision head. The native pool defaults to up to eight available CPUs; override
with JULIA_BEND_THREADS before the first call. JULIA_BEND_TILE_GRAIN overrides
the adaptive projection-specific chunk size. The snapshot rejects changed weights and must be recreated after mutation.
Use the training loader to save or train checkpoints, not an installed backend.
Packed kernels use validated unsafe array sharing and bounded loops; the tree shape proofs do not constitute a formal proof of packed-array memory safety or floating-point math.
Larger choice sets
from julia.router import Router, FastEngine
router = Router(FastEngine('/path/to/checkpoint', device='cuda'), survivors=2)
result = router.route(row_with_up_to_4096_options)
Julia's trained head still supports 2–20 options. Larger choice requests
use batched groups and rerank survivors until a final group remains. A group
retains only its winner when raw softmax gives it over 95% and every other
option is below 4.5%; otherwise it retains the configured survivor count.
This can reduce later model calls for decisive groups, but group probabilities
are not comparable across different groups. This adds model calls and can discard the correct candidate;
it is a capacity feature, not a speed or quality guarantee. Final probabilities
are conditional on result.candidates, never a fabricated global distribution.
model_rows and cache_hits report aggregate work for the whole route_many
call. Optional cache_size caches logits; it defaults to zero. Clear caches
with clear_cache() after changing weights, tokenization or inference settings.
Tests
python -m unittest discover -s julia/router/tests -v
The tests exercise CPU inference and optional native behavior where available.