02 / DEEP LEARNING SYSTEMS

Fluxion
From a tensor to a Transformer. A deep learning engine built from first principles, with its own gradients, trainable layers, GPT model, and native execution experiments.

The chart is generated from the recorded October 4, 2026 Linux CPU run, not a GPU benchmark.
FLUXION / SYSTEM OVERVIEW
The problem
Understanding model training means understanding the machinery underneath it: the computation graph, local derivatives, broadcasting rules, and the costs that remain after moving an operator into native code.
Engineering decisions
- Implemented NumPy-backed tensors, a dynamic DAG, reverse-mode autograd, and shared-input gradient accumulation.
- Composed neural-network layers, losses, SGD/Adam, causal attention, and a small GPT model.
- Checked both outputs and gradients against independent PyTorch implementations.
- Built a portable C++/BLAS Linear operator and experimental standalone CUDA kernels; retained full-workload timing samples.
What the evidence shows
72 CPU/native tests passed and all five PyTorch reference reports passed. Fresh measurements retain 2,000 timed samples. Native Linear was slower on the measured Linux host; different BLAS libraries and framework costs matter. CUDA was not exercised in that checkpoint.
A gradient you can follow
