Atlas — what we are building

A single-company, multi-purpose agentic operating layer orchestrated by Claude. PDF version · Home

We are building a single-company, multi-purpose agentic operating layer: one orchestration mind ("upper mind", run on Claude) that plans, delegates, measures and verifies the work of specialised agents across four internal systems of a dental practice-and-lab business. It is not a chatbot and not a single app; it is the company's working nervous system.

1. Orchestration project (the core). Claude acts as the orchestrator: it writes machine-readable orders with explicit scope, forbidden actions, turn budgets and a single-token verification gate; headless Claude sessions execute them in the right repository with the right model tier (Opus for judgment-heavy changes, Sonnet for implementation, Haiku for high-volume measurement and summaries); results come back as one-line reports that the orchestrator reads before deciding the next step. Scheduled runners (hourly ticks, pre-market and end-of-day digests, nightly audits) run outside any chat session, so the system keeps working when nobody is watching. Every tool carries a self-test (180 tools today), every change passes a measurable gate (dry run over recorded cases, before/after metrics), and every claim is written as a falsifiable hypothesis with a stated "falls if" test.

2. Autonomous screen hand-over. Agents on three machines (a Mac, a lab PC with ExoCAD and Blender, a home PC) hand work to each other through shared inboxes and a task bus: the orchestrator writes a case brief into the lab PC's inbox, the lab agent prepares scans and photos, the Mac agent runs measurements and image generation, results and acceptance records flow back. A clinician never sees the plumbing; they see a photo, a short list of options and a before/after slider on their phone.

3. Dental smile-design pipeline. From two photos (smile, retracted) a local measurement chain extracts tooth proportions, midline, incisal display and gingival margins; patient data never leaves the device. A form-recommendation engine ranks a 441-tooth ExoCAD library using crown geometry read directly from the library meshes (we recovered the cervical line and contact points from the vendor's binary mesh labels), patient profile and clinician preferences; it returns 5 best-fit + 3 exploration slots with written reasons and family-diversity constraints. The chosen form is placed on the photo as a guide (outline and ridge lines) and an image model produces the mockup; accepted designs become ExoCAD work packages.

4. Company management and live-data strategy. Tens of thousands of lines of code read the practice's management system and bank statements, reconcile them, detect anomalies (revenue leaks, scheduling gaps, data-integrity failures) and produce strategic reports for the owner: 150+ reports to date, each scored and with its own "what would change the recommendation" section. Reports are reviewed by a multi-model panel before delivery.

5. Agent-improvement algorithms. The agent family itself is under an explicit evolution loop: hypotheses are logged, tested at the next tick, marked standing or refuted; candidate rule changes are generated from refuted hypotheses, judged by a blind evaluator that cannot see the proposer, measured against placebo changes with Wilson intervals and minimum-detectable-effect budgets, and only consolidated into live rules after a human-approved weekly package. Permutation tests, parrot-detection (copy rate of consecutive judgments) and correlated-agent checks guard against self-reinforcing loops.

6. Theory, thesis and hypothesis rounds ("spark" rounds). For hard design questions we run multi-voice rounds across models that cannot hear each other (Claude, Gemini, Grok, a local Gemma, and a data-seat agent that answers only with repository measurements). Each voice must produce thesis → mechanism → verifiable source → counter-thesis → falsifier. The orchestrator verifies every citation on the web, writes a synthesis that separates agreement from disagreement, records its own prior expectation beforehand and scores afterwards whether the round changed the decision (33 syntheses so far). Several rounds ended with the data seat refuting the clever proposals of the other voices, which is exactly the point.

Scale today. Four repositories, roughly 650,000 lines of Python, 180 self-testing tools, scheduled runners on three machines, 33 spark-round syntheses, 150+ strategic reports, and a working dental pipeline that produced accepted patient mockups this week.

What we need from the program. API credits to move the headless orchestration (nightly measurement runs, multi-agent rounds, mesh analysis over the whole library, Haiku-tier bulk summarisation) from a consumer plan to the API with prompt caching; and engineering guidance on long-running agent workflows, verification gates for autonomous code changes, and cost-aware routing across Opus/Sonnet/Haiku.