| # Julia router |
|
|
| Resident inference for Julia on **CPU or CUDA**, located in the existing `julia` |
| package. No separate repository. Python loads Bend-generated C with `ctypes`; |
| there is no compiler or subprocess in the request path. |
|
|
| ## Build and use |
|
|
| From the **Julia-1** directory (the outer workspace also has an older package |
| named `julia`, so running there can import the wrong one): |
|
|
| ```bash |
| python -m pip install -e . |
| python -m julia.router.build --bend /home/klee/.bend/bin/bend |
| ``` |
|
|
| The native build requires Bend **2.0.27**, Clang, and Linux. It checks `native/PROOF.bend`, generated |
| function arities, and the borrowed-tree contract before compiling. `--native-cpu` |
| adds host-specific machine instructions; use that build only on compatible CPUs. The bridge |
| uses private Bend runtime internals, so upgrading Bend requires adapting and |
| retesting the bridge. `--output` selects another library location; pass that |
| path as `library=` or set `JULIA_ROUTER_LIBRARY` at runtime. |
|
|
| ```python |
| from julia.router import FastEngine |
| |
| engine = FastEngine('/path/to/real/checkpoint', device='cpu', batch_size=16) |
| # device='cuda' uses CUDA/BF16 and device-side softmax/argmax. |
| rows = [{ |
| 'state': 'Preciso trocar minha senha.', |
| 'question': 'Qual é a intenção?', |
| 'options': ['Redefinir senha', 'Cancelar conta', 'Consultar saldo'], |
| }] |
| print(engine.predict(rows)) |
| print(engine.predict(rows, probabilities=False)) # return only indices |
| ``` |
|
|
| Use real checkpoint files, not Git LFS pointer files. The 2026-09-23 runtime audit |
| loaded the actual 551 MiB checkpoint offline and exercised resident inference. |
|
|
| ## Faster inference path |
|
|
| - Bounded LRU caches reuse token fragments and full request encodings. Model |
| forward passes still run on **every request**; no answer cache inflates timings. |
| - Sort by encoded length before bounded microbatches, then restore request order. |
| This reduces padding on mixed-length inputs; it does not guarantee a speedup |
| for every batch or request size. |
| - Build one NumPy arena for integer inputs, with zero-copy CPU tensor views. |
| CUDA uses two pinned bulk transfers instead of per-field transfers. |
| - Softmax and argmax run on the model's device. CUDA does not copy hidden states |
| or attention matrices into CPU Bend kernels. |
| - `compile_model=True` enables optional `torch.compile(dynamic=True)`. Compilation |
| adds first-call cost and requires validation on the target device. |
| - CPU inference retains private file-backed safetensors storage instead of copying |
| the full vocabulary embedding into anonymous RAM. Keep checkpoint files immutable |
| while loaded; use `memory_map=False` for a detached copy. Training model loading |
| still copies by default. |
| - CPU PyTorch heads compute only option queries and feed-forward outputs in the |
| final layer. Full context keys/values are retained. `marker_only_head=False` |
| selects the original path; training and the separate experimental Bend normalization head use the original |
| path automatically. CUDA retains the original default until hardware validation. |
| - `strict_encoding=True` rejects marker injection and any question/option/state |
| truncation. `encoding_info(rows)` audits the same cached encoding used in inference; |
| the game worker no longer tokenizes each request twice. |
| - Preserve Julia's original marker serialization, truncation, weights, and |
| option order. Returned probabilities use the model card presentation rules; |
| `logits()` retains raw scores. |
|
|
| CPU defaults to Python/PyTorch (`torch`), with no Bend build required. CUDA also |
| uses PyTorch. Select `transformer_backend='bend-dense'` explicitly to use the |
| optional CPU FP32 Bend encoder. `compile_model=True` works with the default Torch |
| backend. |
|
|
| ## Bend transformer operations |
|
|
| `native/router.bend` implements candidate argmax, numerically stabilized softmax, |
| and two-pass LayerNorm (mean, centered variance, affine scale/bias). Balanced |
| `Leaf`/`Fork` trees expose candidate reductions. LayerNorm forks over independent |
| rows and uses flat tail loops over feature lists within each row, following the |
| Bend guide's coarse-work/flat-leaf cost model. All use Bend's own heap. The C adapter transports arrays and owns the ABI; it does |
| not duplicate the numerical algorithms. |
|
|
| ```python |
| from julia.router import BendReducer, FastEngine |
| |
| bend = BendReducer() |
| index, probabilities = bend.softmax([1.0, 3.0, -2.0]) |
| normalized = bend.layernorm([[1.0, 2.0, 3.0]]) |
| |
| # Explicit experimental CPU transformer-head backend: |
| engine = FastEngine('/path/to/checkpoint', device='cpu', |
| transformer_backend='bend', bend_postprocess=True) |
| |
| # All 88 encoder projections of the real checkpoint also execute in Bend: |
| # Set JULIA_BEND_THREADS=4 before creating the engine. |
| engine = FastEngine('/path/to/checkpoint', device='cpu', |
| transformer_backend='bend-dense', bend_postprocess=True) |
| ``` |
|
|
| The experimental head executes its pre-attention and pre-MLP LayerNorms, plus |
| scorer LayerNorm, in Bend. Dense projections, attention, and the encoder remain |
| PyTorch. This is **not a complete transformer rewrite in Bend**. Bend head math |
| is inference-only CPU float32. The CPU library defaults to up to eight available Bend runtime |
| workers. Set `JULIA_BEND_THREADS=4` **before the first native call** to test row |
| parallelism; the worker count is fixed for the process. Pool sizing depends on the kernel and workload; the older normalization-only |
| experiment does not determine the dense backend defaults. Reduction order differs from the |
| old tree algorithm and PyTorch: probabilities can differ by floating-point rounding, |
| and effectively tied choices may choose a different index. |
| Calls share a mutex because Bend's runtime has global state. Instantiate model |
| workers in spawned processes; do not fork an active inference process. |
|
|
| ### Resident dense projections (CPU default) |
|
|
| `bend-dense` executes all 88 encoder projections in Bend. Resident packed FP32 |
| arrays replace linked weight trees. Contiguous 8×8 tiles expose vector arithmetic; |
| coarse parallel ranges end in flat tail loops. Shared input/weight handles are |
| read-only, output tiles are disjoint, and all handles are joined after evaluation. |
| The bridge packs/transports buffers and checks bounds; the public API validates |
| finite values, while engine inference validates weights once and final logits. |
|
|
| PyTorch retains embeddings, SDPA, elementwise activations and the selected-output |
| decision head. The native pool defaults to up to eight available CPUs; override |
| with `JULIA_BEND_THREADS` before the first call. `JULIA_BEND_TILE_GRAIN` overrides |
| the adaptive projection-specific chunk size. The snapshot rejects changed weights and must be recreated after mutation. |
| Use the training loader to save or train checkpoints, not an installed backend. |
|
|
| Packed kernels use |
| validated unsafe array sharing and bounded loops; the tree shape proofs do not |
| constitute a formal proof of packed-array memory safety or floating-point math. |
|
|
| ## Larger choice sets |
|
|
| ```python |
| from julia.router import Router, FastEngine |
| router = Router(FastEngine('/path/to/checkpoint', device='cuda'), survivors=2) |
| result = router.route(row_with_up_to_4096_options) |
| ``` |
|
|
| Julia's trained head still supports **2–20** options. Larger *choice* requests |
| use batched groups and rerank survivors until a final group remains. A group |
| retains only its winner when raw softmax gives it over 95% and every other |
| option is below 4.5%; otherwise it retains the configured survivor count. |
| This can reduce later model calls for decisive groups, but group probabilities |
| are not comparable across different groups. This adds model calls and can discard the correct candidate; |
| it is a capacity feature, not a speed or quality guarantee. Final probabilities |
| are conditional on `result.candidates`, never a fabricated global distribution. |
| `model_rows` and `cache_hits` report aggregate work for the whole `route_many` |
| call. Optional `cache_size` caches logits; it defaults to zero. Clear caches |
| with `clear_cache()` after changing weights, tokenization or inference settings. |
|
|
| ## Tests |
|
|
| ```bash |
| python -m unittest discover -s julia/router/tests -v |
| ``` |
|
|
| The tests exercise CPU inference and optional native behavior where available. |
|
|