Julia-1 / julia /router /README.md
kleeedolinux
Publish Julia 1 model and Python runtime
5278d6b
|
Raw
History Blame Contribute Delete
8.21 kB
# Julia router
Resident inference for Julia on **CPU or CUDA**, located in the existing `julia`
package. No separate repository. Python loads Bend-generated C with `ctypes`;
there is no compiler or subprocess in the request path.
## Build and use
From the **Julia-1** directory (the outer workspace also has an older package
named `julia`, so running there can import the wrong one):
```bash
python -m pip install -e .
python -m julia.router.build --bend /home/klee/.bend/bin/bend
```
The native build requires Bend **2.0.27**, Clang, and Linux. It checks `native/PROOF.bend`, generated
function arities, and the borrowed-tree contract before compiling. `--native-cpu`
adds host-specific machine instructions; use that build only on compatible CPUs. The bridge
uses private Bend runtime internals, so upgrading Bend requires adapting and
retesting the bridge. `--output` selects another library location; pass that
path as `library=` or set `JULIA_ROUTER_LIBRARY` at runtime.
```python
from julia.router import FastEngine
engine = FastEngine('/path/to/real/checkpoint', device='cpu', batch_size=16)
# device='cuda' uses CUDA/BF16 and device-side softmax/argmax.
rows = [{
'state': 'Preciso trocar minha senha.',
'question': 'Qual é a intenção?',
'options': ['Redefinir senha', 'Cancelar conta', 'Consultar saldo'],
}]
print(engine.predict(rows))
print(engine.predict(rows, probabilities=False)) # return only indices
```
Use real checkpoint files, not Git LFS pointer files. The 2026-09-23 runtime audit
loaded the actual 551 MiB checkpoint offline and exercised resident inference.
## Faster inference path
- Bounded LRU caches reuse token fragments and full request encodings. Model
forward passes still run on **every request**; no answer cache inflates timings.
- Sort by encoded length before bounded microbatches, then restore request order.
This reduces padding on mixed-length inputs; it does not guarantee a speedup
for every batch or request size.
- Build one NumPy arena for integer inputs, with zero-copy CPU tensor views.
CUDA uses two pinned bulk transfers instead of per-field transfers.
- Softmax and argmax run on the model's device. CUDA does not copy hidden states
or attention matrices into CPU Bend kernels.
- `compile_model=True` enables optional `torch.compile(dynamic=True)`. Compilation
adds first-call cost and requires validation on the target device.
- CPU inference retains private file-backed safetensors storage instead of copying
the full vocabulary embedding into anonymous RAM. Keep checkpoint files immutable
while loaded; use `memory_map=False` for a detached copy. Training model loading
still copies by default.
- CPU PyTorch heads compute only option queries and feed-forward outputs in the
final layer. Full context keys/values are retained. `marker_only_head=False`
selects the original path; training and the separate experimental Bend normalization head use the original
path automatically. CUDA retains the original default until hardware validation.
- `strict_encoding=True` rejects marker injection and any question/option/state
truncation. `encoding_info(rows)` audits the same cached encoding used in inference;
the game worker no longer tokenizes each request twice.
- Preserve Julia's original marker serialization, truncation, weights, and
option order. Returned probabilities use the model card presentation rules;
`logits()` retains raw scores.
CPU defaults to Python/PyTorch (`torch`), with no Bend build required. CUDA also
uses PyTorch. Select `transformer_backend='bend-dense'` explicitly to use the
optional CPU FP32 Bend encoder. `compile_model=True` works with the default Torch
backend.
## Bend transformer operations
`native/router.bend` implements candidate argmax, numerically stabilized softmax,
and two-pass LayerNorm (mean, centered variance, affine scale/bias). Balanced
`Leaf`/`Fork` trees expose candidate reductions. LayerNorm forks over independent
rows and uses flat tail loops over feature lists within each row, following the
Bend guide's coarse-work/flat-leaf cost model. All use Bend's own heap. The C adapter transports arrays and owns the ABI; it does
not duplicate the numerical algorithms.
```python
from julia.router import BendReducer, FastEngine
bend = BendReducer()
index, probabilities = bend.softmax([1.0, 3.0, -2.0])
normalized = bend.layernorm([[1.0, 2.0, 3.0]])
# Explicit experimental CPU transformer-head backend:
engine = FastEngine('/path/to/checkpoint', device='cpu',
transformer_backend='bend', bend_postprocess=True)
# All 88 encoder projections of the real checkpoint also execute in Bend:
# Set JULIA_BEND_THREADS=4 before creating the engine.
engine = FastEngine('/path/to/checkpoint', device='cpu',
transformer_backend='bend-dense', bend_postprocess=True)
```
The experimental head executes its pre-attention and pre-MLP LayerNorms, plus
scorer LayerNorm, in Bend. Dense projections, attention, and the encoder remain
PyTorch. This is **not a complete transformer rewrite in Bend**. Bend head math
is inference-only CPU float32. The CPU library defaults to up to eight available Bend runtime
workers. Set `JULIA_BEND_THREADS=4` **before the first native call** to test row
parallelism; the worker count is fixed for the process. Pool sizing depends on the kernel and workload; the older normalization-only
experiment does not determine the dense backend defaults. Reduction order differs from the
old tree algorithm and PyTorch: probabilities can differ by floating-point rounding,
and effectively tied choices may choose a different index.
Calls share a mutex because Bend's runtime has global state. Instantiate model
workers in spawned processes; do not fork an active inference process.
### Resident dense projections (CPU default)
`bend-dense` executes all 88 encoder projections in Bend. Resident packed FP32
arrays replace linked weight trees. Contiguous 8×8 tiles expose vector arithmetic;
coarse parallel ranges end in flat tail loops. Shared input/weight handles are
read-only, output tiles are disjoint, and all handles are joined after evaluation.
The bridge packs/transports buffers and checks bounds; the public API validates
finite values, while engine inference validates weights once and final logits.
PyTorch retains embeddings, SDPA, elementwise activations and the selected-output
decision head. The native pool defaults to up to eight available CPUs; override
with `JULIA_BEND_THREADS` before the first call. `JULIA_BEND_TILE_GRAIN` overrides
the adaptive projection-specific chunk size. The snapshot rejects changed weights and must be recreated after mutation.
Use the training loader to save or train checkpoints, not an installed backend.
Packed kernels use
validated unsafe array sharing and bounded loops; the tree shape proofs do not
constitute a formal proof of packed-array memory safety or floating-point math.
## Larger choice sets
```python
from julia.router import Router, FastEngine
router = Router(FastEngine('/path/to/checkpoint', device='cuda'), survivors=2)
result = router.route(row_with_up_to_4096_options)
```
Julia's trained head still supports **2–20** options. Larger *choice* requests
use batched groups and rerank survivors until a final group remains. A group
retains only its winner when raw softmax gives it over 95% and every other
option is below 4.5%; otherwise it retains the configured survivor count.
This can reduce later model calls for decisive groups, but group probabilities
are not comparable across different groups. This adds model calls and can discard the correct candidate;
it is a capacity feature, not a speed or quality guarantee. Final probabilities
are conditional on `result.candidates`, never a fabricated global distribution.
`model_rows` and `cache_hits` report aggregate work for the whole `route_many`
call. Optional `cache_size` caches logits; it defaults to zero. Clear caches
with `clear_cache()` after changing weights, tokenization or inference settings.
## Tests
```bash
python -m unittest discover -s julia/router/tests -v
```
The tests exercise CPU inference and optional native behavior where available.