← Cambrian Weekly Blog
Houston · Weekly Update

Project Houston Weekly · June 22–26 (Week 1)

Achuta KadambiAchuta Kadambi Friday, June 26, 2026 ·4 min read

Houston is the brain that coordinates our jobsites. It’s named for being the “mission control” of a jobsite. Each jobsite can carry thousands of tasks spread across hundreds of vendors, generating hundreds of invoices.

Software to track this stuff is known as “Enterprise Resource Planning” (ERP), used by business software organizations to integrate, automate, and manage their core operations.

Five years ago, ERPs were for larger companies and took many engineering hours to install in an org: the average ERP for mid-market companies took $4.5M to install over 14 months, and 77% of companies brought in outside consultants during the setup (Panorama Consulting).

LLMs make it much easier to build an ERP today, but there are several challenges. The simplest ERP is just an LLM running over business files, chats, and emails. Unfortunately, an LLM will start to hallucinate and take wrong actions when it is overwhelmed by so much data. This is because there are a lot of irrelevant files that will get stuck in the context window of the LLM.

To solve this problem, Houston is a “semi-deterministic” ERP. As shown in the figure, we define a strict schema for how a request gets handled: the LLM interprets it but does not take the action itself; it routes it through a deterministic sequence of methods, each with a defined scope of what it’s allowed to touch.

Houston converts natural language into ERP functions.
Houston converts natural language into ERP functions.

What got built, last week to this week

Last week we set up the APIs and the early file schema to establish a “source of truth”. The source of truth establishes which documents the company treats as authoritative, so we don’t get duplicates floating around. Then we set Houston loose with a raw LLM API just to see how it would fare. Although we were happy the API integrations were working, the performance of Houston did not go well:

Before this week

Houston · Marinette
AK
Achuta9:14 AM
@Houston the 2725 comp link looks wrong
H
HoustonAPP9:14 AM
I need project_id, property — which?

Here, the problem is that the LLM had trouble obtaining the context from the prompt (it didn’t know that 2725 referred to 2725 Midvale avenue which was part of a separate meeting between Achuta and Sasha). We fixed that this week with a few commits:

After applying these commits, we returned the following:

After this week

Houston · Marinette
AK
Achuta9:14 AM
@Houston the 2725 comp link looks wrong
H
HoustonAPP9:14 AM
Got it — flagged the 2725 comp for review. No changes made.

And if one strengthens the intent to become more “command-like”, Houston can take the action with scoped permissions:

After this week

Houston · Marinette
AK
Achuta9:18 AM
@Houston add this as a Marinette comp redfin.com/…/2725
H
HoustonAPP9:18 AM
Added it as a Marinette comp — pulled the sold price & days-on-market from the listing, deduped against existing comps, logged it.

What do you mean by “semi-deterministic”?

It means we use the LLM to route a request to a deterministic function that does a defined thing like create, read, update, or delete a file (CRUD).

What eval did we run? (numbers checked against the eval report)

We use a separate agent to generate adversarial test cases with labeled ground truth, then run two evals.

Is 96% good enough? Not quite, but it’s not an issue with the LLM routing. That last ~4% is actually due to an ID collision in the registry cache, which can be fixed.

Prompt injection testing

If an employee with good intentions writes the wrong message to Houston, could it delete all of Cambrian’s files? Because the real work happens in deterministic functions, we can scope them explicitly. We ran 25 prompt-injection attempts, which are messages trying to make Houston leak secrets or fire bad writes. These harmful messages returned 0 leaks and 0 malicious routes, which is a good start (though further testing is needed).

Why not just buy Procore or Oracle, or use the new Claude Tag?

We could buy an ERP, and there are some good options out there. The decision to build our own is less about pricing and more about what’s right for our model.

Procore, Buildertrend, Oracle, NetSuite etc. can be set up in a way that it can really help traditional homebuilding workflows, but there are a couple reasons why these solutions may not be right for Cambrian. A first reason is that existing construction ERP workflows are geared around known templates while custom homebuilding can change house to house.

A second, and perhaps more significant, reason is that the sci-fi technology does not play nicely with containerized ERP workflows. For example, we want to use world model & computer vision feedback loops to change the ERP task pipelining. That’s hard to integrate with existing ERP APIs today. This will be even more true as Cambrian starts to become more ambitious in bringing in new tasks like 3D scanning, site observation, humanoid robot workers, and so on.

With AI models getting so good, an emerging solution is to use an Enterprise AI from a model provider like Claude on your data. This works, but if it doesn’t have guardrails the hallucinations can cause chaos on our processes. Application-layer companies in AI that sit on top like Glean do offer solutions, but then these solutions are more geared toward traditional business intelligence workflows (sales, marketing, etc), not homebuilding.

This is a story that played out at SpaceX many years ago, where the specific workflows they needed to build rockets efficiently were not being covered by legacy providers. So they buit WarpDrive internally as an alternative. If it made sense for SpaceX to build it from scratch in an era where software was very expensive, it seems like we should build it internally today, in this era where software has never been easier to build.

How can I try Houston?

Join the Houston space in G-Suite and tag it with a question — like “@Houston tell me why Cambrian is a cool company.” It only acts when you actually ask it to.

What’s the goal for next week?

A few things:

Houston backlog

Item What it is Status
Houston → cloud Move the always-on loop off the laptop to a cloud host — no single point of failure 💡 idea
Self-improving Houston It proposes its own changes (scoped diff + tests + 👍), never self-merges 💡 idea
Skills in Google Chat A discoverable slash-skill catalog (/comps, /vendor, /ticket, /status…) for conversational ERP work 💡 idea
Secrets & auth hardening Credentials off local disk into a secret store; a service account, not a personal login 💡 idea
Observability One queryable log of every action, the reasoning, and the approval that gated it 💡 idea
Continuous eval A regression suite that runs on every change, not point-in-time spot checks 💡 idea
Multi-project scale-out “Add a project” as one command/chat; fan out cleanly across N projects 💡 idea
Cost & usage tracking Token + dollar usage per project per run, before it’s a surprise 💡 idea
Retire nightly jobs Move recordings/calendar to always-on workers; drop the nightly cron 💡 idea
Vendor registry Structured, queryable vendor records — harvested at ingest, identity-resolved, evidence-cited human scoring 🔎 scoped