aikaara
← notes

2026-08-11T00:00:00.000Z

Installing the factory: what on-prem agents actually need

When an enterprise asks to run the factory inside their walls, the interesting questions are not about the agents. They are about the environment the agents live in. Here is what on-prem agents actually need, from the pilot outward.

Models you host, benchmarked before you trust them

Frontier capability lives in hosted models you cannot bring inside an air gap. On-prem means open-weight models you run yourself, and they do not all handle every workload equally. So the first deliverable is not an agent — it is a benchmark on your hardware: which of your candidate models handles which workloads, measured, not asserted. You decide what to trust from numbers, not vendor claims.

An evidence trail, treated as a product

Inside a regulated enterprise, “the agent did it” is not an answer. Every agent action has to be logged as a first-class artifact: what was built, what changed, what was deployed, by which agent, approved by whom. This evidence pack is not a debug log you keep in case of trouble — it is a deliverable the pilot produces on purpose, because sovereignty without accountability is not sovereignty.

A named human on the outcome

Throughput comes from agents. Ownership comes from a person. The factory does the volume; a named operator signs off on what ships and answers for it. The org chart does not disappear — it gets one clear line of accountability instead of a diffuse “the AI.”

Honest hardware math, up front

On-prem means GPUs, and GPUs mean procurement. We size the hardware in the first meeting, in writing, before anything else — no surprise capital request three weeks in. The pilot is fixed price, two to four weeks, one shipped workload, and a go / no-go recommendation you can take to a budget owner.

The factory is not magic that ignores your constraints. It is a system that respects them and produces evidence at every step. That is the whole point of installing it rather than renting an outcome.