Local LLM Engineering Notes
Technical notes and reproducible observations from redoracle while exploring local LLM inference, GGUF artifacts, llama.cpp and AMD ROCm systems.
This repository is a durable index for engineering work. It records measured behavior with its hardware, runtime, model identity, configuration and limitations. It does not present universal performance rankings or unsupported compatibility claims.
Projects
- GGUF Inspector parses GGUF headers, metadata and tensor descriptors locally in the browser. It does not execute models or validate numerical correctness.
- Local LLM Systems Lab records controlled local inference experiments with versioned machine readable results.
Case study
The Nemotron GGUF study documents a reproducible metadata consistency failure, its isolation through a metadata only diagnostic copy, CPU and ROCm controls, and the eventual upstream resolution in llama.cpp PR 27729.
The diagnostic metadata edit was a causality test, not a recommended artifact repair procedure. Runtime, backend and model quality conclusions remain scoped to the recorded tests.
Working principles
- Inspect artifact identity and metadata before attributing a failure to hardware.
- Separate parsing, structural checks, architecture checks and runtime validation.
- Record unknowns and limitations instead of inferring missing evidence.
- Prefer upstream fixes and reproducible reports over local workarounds.
- Keep experiments small, reviewable and useful to other engineers.
Scope
The work focuses on local inference systems, GGUF format behavior, quantization workflows, llama.cpp runtime diagnostics, AMD ROCm observations and practical AI tooling. Results are educational engineering records, not formal research claims.