I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P]

Reddit r/MachineLearning Tools

Summary

A compiler that converts computation graphs defined in Python into the weights of a standard transformer architecture (Phi-3) without any training, enabling vanilla HuggingFace deployment with no custom code.

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it produces the weights of a transformer that executes the graph. The result is a standard Phi-3-architecture checkpoint that vanilla huggingface loads with no custom code and no trust_remote_code. Zero training in the pipeline. Write-up (origin + how the constructions work): https://ood.dev/posts/torchwright-intro/ Repo (twelve runnable examples): https://github.com/physicsrob/torchwright Hand-built transformer weights aren't a new idea. RASP defines a language whose primitives map onto transformer sublayers, and Tracr compiles RASP programs into actual weights. I wanted two things they don't aim for: expressing a computation graph in ordinary Python, and targeting a stock architecture, so the output loads in vanilla huggingface with no custom code.
Original Article

Similar Articles

Transformers now runs llama.cpp quants

Hugging Face Blog

Hugging Face's transformers library now supports GGUF models from llama.cpp, enabling efficient local inference on consumer hardware through familiar APIs.

Transformer Math Explorer [P]

Reddit r/MachineLearning

This interactive tool visualizes the mathematical underpinnings of transformer models through dataflow graphs, covering architectures from GPT-2 to Qwen 3.6 and various attention mechanisms.