Julia-1 / julia /router /README.md
kleeedolinux
Publish Julia 1 model and Python runtime
5278d6b
|
Raw
History Blame Contribute Delete
8.21 kB

Julia router

Resident inference for Julia on CPU or CUDA, located in the existing julia package. No separate repository. Python loads Bend-generated C with ctypes; there is no compiler or subprocess in the request path.

Build and use

From the Julia-1 directory (the outer workspace also has an older package named julia, so running there can import the wrong one):

python -m pip install -e .
python -m julia.router.build --bend /home/klee/.bend/bin/bend

The native build requires Bend 2.0.27, Clang, and Linux. It checks native/PROOF.bend, generated function arities, and the borrowed-tree contract before compiling. --native-cpu adds host-specific machine instructions; use that build only on compatible CPUs. The bridge uses private Bend runtime internals, so upgrading Bend requires adapting and retesting the bridge. --output selects another library location; pass that path as library= or set JULIA_ROUTER_LIBRARY at runtime.

from julia.router import FastEngine

engine = FastEngine('/path/to/real/checkpoint', device='cpu', batch_size=16)
# device='cuda' uses CUDA/BF16 and device-side softmax/argmax.
rows = [{
    'state': 'Preciso trocar minha senha.',
    'question': 'Qual é a intenção?',
    'options': ['Redefinir senha', 'Cancelar conta', 'Consultar saldo'],
}]
print(engine.predict(rows))
print(engine.predict(rows, probabilities=False))  # return only indices

Use real checkpoint files, not Git LFS pointer files. The 2026-09-23 runtime audit loaded the actual 551 MiB checkpoint offline and exercised resident inference.

Faster inference path

  • Bounded LRU caches reuse token fragments and full request encodings. Model forward passes still run on every request; no answer cache inflates timings.
  • Sort by encoded length before bounded microbatches, then restore request order. This reduces padding on mixed-length inputs; it does not guarantee a speedup for every batch or request size.
  • Build one NumPy arena for integer inputs, with zero-copy CPU tensor views. CUDA uses two pinned bulk transfers instead of per-field transfers.
  • Softmax and argmax run on the model's device. CUDA does not copy hidden states or attention matrices into CPU Bend kernels.
  • compile_model=True enables optional torch.compile(dynamic=True). Compilation adds first-call cost and requires validation on the target device.
  • CPU inference retains private file-backed safetensors storage instead of copying the full vocabulary embedding into anonymous RAM. Keep checkpoint files immutable while loaded; use memory_map=False for a detached copy. Training model loading still copies by default.
  • CPU PyTorch heads compute only option queries and feed-forward outputs in the final layer. Full context keys/values are retained. marker_only_head=False selects the original path; training and the separate experimental Bend normalization head use the original path automatically. CUDA retains the original default until hardware validation.
  • strict_encoding=True rejects marker injection and any question/option/state truncation. encoding_info(rows) audits the same cached encoding used in inference; the game worker no longer tokenizes each request twice.
  • Preserve Julia's original marker serialization, truncation, weights, and option order. Returned probabilities use the model card presentation rules; logits() retains raw scores.

CPU defaults to Python/PyTorch (torch), with no Bend build required. CUDA also uses PyTorch. Select transformer_backend='bend-dense' explicitly to use the optional CPU FP32 Bend encoder. compile_model=True works with the default Torch backend.

Bend transformer operations

native/router.bend implements candidate argmax, numerically stabilized softmax, and two-pass LayerNorm (mean, centered variance, affine scale/bias). Balanced Leaf/Fork trees expose candidate reductions. LayerNorm forks over independent rows and uses flat tail loops over feature lists within each row, following the Bend guide's coarse-work/flat-leaf cost model. All use Bend's own heap. The C adapter transports arrays and owns the ABI; it does not duplicate the numerical algorithms.

from julia.router import BendReducer, FastEngine

bend = BendReducer()
index, probabilities = bend.softmax([1.0, 3.0, -2.0])
normalized = bend.layernorm([[1.0, 2.0, 3.0]])

# Explicit experimental CPU transformer-head backend:
engine = FastEngine('/path/to/checkpoint', device='cpu',
                    transformer_backend='bend', bend_postprocess=True)

# All 88 encoder projections of the real checkpoint also execute in Bend:
# Set JULIA_BEND_THREADS=4 before creating the engine.
engine = FastEngine('/path/to/checkpoint', device='cpu',
                    transformer_backend='bend-dense', bend_postprocess=True)

The experimental head executes its pre-attention and pre-MLP LayerNorms, plus scorer LayerNorm, in Bend. Dense projections, attention, and the encoder remain PyTorch. This is not a complete transformer rewrite in Bend. Bend head math is inference-only CPU float32. The CPU library defaults to up to eight available Bend runtime workers. Set JULIA_BEND_THREADS=4 before the first native call to test row parallelism; the worker count is fixed for the process. Pool sizing depends on the kernel and workload; the older normalization-only experiment does not determine the dense backend defaults. Reduction order differs from the old tree algorithm and PyTorch: probabilities can differ by floating-point rounding, and effectively tied choices may choose a different index. Calls share a mutex because Bend's runtime has global state. Instantiate model workers in spawned processes; do not fork an active inference process.

Resident dense projections (CPU default)

bend-dense executes all 88 encoder projections in Bend. Resident packed FP32 arrays replace linked weight trees. Contiguous 8×8 tiles expose vector arithmetic; coarse parallel ranges end in flat tail loops. Shared input/weight handles are read-only, output tiles are disjoint, and all handles are joined after evaluation. The bridge packs/transports buffers and checks bounds; the public API validates finite values, while engine inference validates weights once and final logits.

PyTorch retains embeddings, SDPA, elementwise activations and the selected-output decision head. The native pool defaults to up to eight available CPUs; override with JULIA_BEND_THREADS before the first call. JULIA_BEND_TILE_GRAIN overrides the adaptive projection-specific chunk size. The snapshot rejects changed weights and must be recreated after mutation. Use the training loader to save or train checkpoints, not an installed backend.

Packed kernels use validated unsafe array sharing and bounded loops; the tree shape proofs do not constitute a formal proof of packed-array memory safety or floating-point math.

Larger choice sets

from julia.router import Router, FastEngine
router = Router(FastEngine('/path/to/checkpoint', device='cuda'), survivors=2)
result = router.route(row_with_up_to_4096_options)

Julia's trained head still supports 2–20 options. Larger choice requests use batched groups and rerank survivors until a final group remains. A group retains only its winner when raw softmax gives it over 95% and every other option is below 4.5%; otherwise it retains the configured survivor count. This can reduce later model calls for decisive groups, but group probabilities are not comparable across different groups. This adds model calls and can discard the correct candidate; it is a capacity feature, not a speed or quality guarantee. Final probabilities are conditional on result.candidates, never a fabricated global distribution. model_rows and cache_hits report aggregate work for the whole route_many call. Optional cache_size caches logits; it defaults to zero. Clear caches with clear_cache() after changing weights, tokenization or inference settings.

Tests

python -m unittest discover -s julia/router/tests -v

The tests exercise CPU inference and optional native behavior where available.