Selected work / Fluxion

02 / DEEP LEARNING SYSTEMS

Fluxion logo

Fluxion

From a tensor to a Transformer. A deep learning engine built from first principles, with its own gradients, trainable layers, GPT model, and native execution experiments.

Measured Fluxion GPT CPU training-step latency and throughput across sequence lengths

The chart is generated from the recorded October 4, 2026 Linux CPU run, not a GPU benchmark.

FLUXION / SYSTEM OVERVIEW

TensorAutogradLayersGPTIndependent output and gradient checks
PythonNumPyC++

The problem

Understanding model training means understanding the machinery underneath it: the computation graph, local derivatives, broadcasting rules, and the costs that remain after moving an operator into native code.

Engineering decisions

  • Implemented NumPy-backed tensors, a dynamic DAG, reverse-mode autograd, and shared-input gradient accumulation.
  • Composed neural-network layers, losses, SGD/Adam, causal attention, and a small GPT model.
  • Checked both outputs and gradients against independent PyTorch implementations.
  • Built a portable C++/BLAS Linear operator and experimental standalone CUDA kernels; retained full-workload timing samples.

What the evidence shows

72 CPU/native tests passed and all five PyTorch reference reports passed. Fresh measurements retain 2,000 timed samples. Native Linear was slower on the measured Linux host; different BLAS libraries and framework costs matter. CUDA was not exercised in that checkpoint.

A gradient you can follow

Shared-input autograd example: y equals x times x plus x. At x=3, three contributions add to a gradient of 7.