faxl

faxl

navigate the frontier

faxl makes privately hosted sovereign AI easy and cost effective.

Get faxl — free

Why faxl

What is an agent?

Hidden from view, every AI agent is defined by long internal system prompts that get sent along with every question you ask. They define its personality, its rules of thumb, its moral code. Its soul.

Each time you send a query, those prompts are added to the front of yours. Saying "Hi" can burn 80,000 tokens as the model reads all of it first.

These injected personalities are not the problem. They are the answer. Skill files, domain knowledge, compacted history, rules, style, reminders — all of it goes in front of what you write, and that is what makes the experience good and the work consistent.

The cost

Agentic traffic needs a prefill cache

What does all that cost? It depends. If your AI service has a decent prefill cache, the GPU remembers those prefix sequences and starts from a checkpoint you already paid for, instead of charging you twice.

That cache runs by rules: what is kept, for how long, which checkpoints win across machines and customers. Those decisions are not yours. Mostly we would like to believe the providers are doing their best for everyone.

But what if you need 100% private AI, and want to run your own hardware, or bare metal you rent? Who manages that cache now? Nobody. It is yours to organise, and that is what makes the move to sovereign hardware hard. The providers have no reason to share this technology — the lack of options is the moat.

faxl

The missing piece

faxl is the software you need to run your own caching layer, in front of your own models. It can save 90% of local token burn. You manage it as you see fit. It lets you leave the cloud services and go private.

Measured

What it does, on real traffic

20.5× faster first token on a repeated prompt — 5,056 ms to 246 ms
~70% of prompt tokens skipped on real agent traffic through a warm store
570/570 adversarial reuse decisions correct; no wrong or stale state served

The 20.5× figure is Kimi Linear 48B, one prompt class, on Apple silicon. Your result depends on how much your prompts repeat. faxl helps when prompts share long prefixes: system prompts, RAG preambles, document Q&A, agent loops. If yours do not, faxl gives you the analytics to work out how to reorganise your agentic prompts — to engineer them, and take control of your performance.

Correctness

Reuse is verified, not assumed

A fingerprint match is never trusted on its own. faxl confirms the candidate token-for-token before the model resumes from it, so output is byte-identical to an uncached run. A miss costs a little lookup time — never a wrong answer.

How it works →

Today

Apple silicon now, NVIDIA next

faxl runs on a Mac with an M-series chip via MLX — including a Mac Studio you have already bought and are not getting full use of. It also runs across several Macs over Thunderbolt, which is how you run a model too large for any one of them. The NVIDIA/vLLM port passes the same byte-identity bars on GPU.

Pre-launch: the first 1,000 registrations get a free named licence, valid 30 days, for noncommercial use.

Get faxl — free