What happens when somebody gets in touch
Start with one message from one visitor and follow it. Almost everything in this document is visible in that single turn: how the message is handled, what the agent can and cannot reach, when qualification happens, and what has to be true before anything is written down or sent.
The top row is DMA's process as I currently understand it — assembled from the v3 Canonical and our two calls, and reported rather than proposed. I have no reason to think I understand how DMA sells better than DMA does, and the parts I am least sure of are the ones nobody has written down. Where this drawing and your actual practice disagree, your practice is right and the drawing is the thing that is wrong. Correcting it is the first hour of Discovery, and it is cheap now and expensive later.
Everything below the top row is mine — how I would build the machinery so that those stages happen reliably, and the same way every time. So there are two different kinds of disagreement available in these pictures, and both are useful: argue with the top row on the facts, and with everything under it on the engineering.
The point of automating a sales motion is to make what already works happen more often. A system built on a process the team does not recognise is worse than no system, which is why this page comes first.
Three things worth reading off it
The visitor's words and the operating instructions never travel together. The message arrives on the content channel as quoted data; the policy — who the agent is, what it may do — arrives separately and cannot be amended by anything a visitor types. That separation, plus a tool list fixed before the turn starts, is the whole of the containment story: the agent's reach is the seven tools in that box and nothing else, no matter how persuasive the message.
Qualification is not a stage somebody declares. It runs after every single turn. The transcript is re-read, A03 pulls out services, scale, timing, authority and budget signal — each one stored with the sentence it came from — and a scorer written in code, not a model, turns those facts into a number and a tier. Nobody has to notice that a visitor became interesting; the threshold does, on the turn it becomes true.
The fast path and the committing path are deliberately different. A reply is checked for unresolved claims and shown in seconds, because a prospect will not wait. Anything that changes state — a Nutshell write, a hot-lead alert, a booking, a hand-off — becomes an intent and queues for the gate, where it is re-validated against live state before it commits. The visitor is never waiting on the gate, and the gate is never rushed by the visitor.
The same shape, across the whole journey
That loop is not special to first contact. It repeats at every stage: something arrives, an agent proposes, the gate decides, a person sees the result. What changes from stage to stage is which agents are involved and how much the gate is willing to let through on its own. The stages themselves are the part to check against how DMA actually sells — if a step is missing, out of order, or never happens the way it is drawn here, that is worth saying before anything is built on it.
Three things to take from the map
The top row is the only part anyone outside DMA sees. Everything else — fifteen agents, a scoring pipeline, a mirrored database, a nightly reconciliation — exists to make those six boxes happen reliably. That is also why almost none of this system has a user interface, and why this arrived as a set of documents rather than as a demo: the parts that decide whether it works are not the parts a demo can show. Part Two, section 00 makes that case.
The gate is a single line every action crosses, not a step in a sequence. Drafting an e-mail, writing a CRM field, sending a proposal and charging a deposit are wildly different actions, and all four are re-validated, checked against DMA's rules and committed exactly once by the same piece of code. It is the first thing I would build and the thing most likely to be quietly missing from a system that demos well.
The bottom band is where the money is. Before a dormant record ever becomes a first contact, it has been read, scored, tiered and given a reason to be worth contacting today. Thirty thousand records are the fastest path to real revenue, which is why that work comes first in the build sequence rather than the website chat.
What this is, and what it deliberately is not
This is a surface and control map, not a proof of concept. It answers one question in detail: when the Canonical says the AI qualifies, scores, reaches out, negotiates, proposes, signs and collects — where does a human actually look, and what is the software that keeps all of it honest?
A generated proof of concept would have been the easier thing to send, and I would rather not lead with one. It shows a UI that has never touched your Nutshell account, priced against no pricing policy, with no idea whether your API supports the writes it assumes. It looks finished and it is unfalsifiable. It also fails the Canonical, which says plainly: no custom approval UI, Nutshell remains the single system of record.
So this document does the opposite. It takes your requirements as written, follows them to their consequences, and shows you the consequences — including the four or five places where the Canonical asks for two things that cannot both be true. Every screen illustrated here is a sketch of a real surface I would expect to build or configure, drawn to show placement and content, not visual design.
Almost none of this system has a user interface. The reactivation engine, the scoring pass, the reply handling, the write-back — all of it is background work that surfaces inside Nutshell as fields, tasks and lists your reps already know how to read. The custom software you actually need is small, and it is not a sales tool: it is the control plane that holds the guardrails, the action queue, the audit log and the transaction state, and reps will rarely open it.
I have not put a fixed price in this document, and the milestone dates below are shapes rather than commitments. That is a deliberate position rather than an evasion: the two largest cost variables in this project — whether the Nutshell API supports the writes the design assumes, and how much of the 92 Phase 1 rows constitutes the first release — are both unmeasured. Anyone quoting a firm number today is quoting across a gap they have not looked into. The sequence in section 11 is real; the calendar it lands on gets fixed once those two things are known. Section 05 does put numbers on the recurring model and vendor spend, which is a different question and an answerable one.
Two engines, one commit gate, one system of record
The Canonical's architecture page describes "one AI brain." On a technical level that is a useful way to talk about it but not a thing you build. What you build is a large number of small, single-purpose agents — each one a prompt, an input schema and an output schema — plus a shared account memory they all read from, plus one piece of deterministic plumbing that every single one of them has to pass through before it is allowed to change anything in the real world.
That last piece is the whole design. It is what stops fifteen agents from stepping on each other, on your reps, and on Nutshell.
Consider what happens when the AI is mid-negotiation, a rep is editing the same record, and a webhook fires — all at once. The honest answer is that no amount of prompt engineering solves that, because you cannot make fifteen independent language models agree on a locking protocol. It has to be solved once, deterministically, outside the agents. That is what the gate is, and it is the first thing I would build after the Nutshell adapter.
Every Phase 1 human touchpoint, and where it lives
This is the table the Canonical is missing. It takes each Phase 1 capability that a person has to see, touch or approve, and assigns it a venue. Four venues only: Nutshell, a channel you already read, a product you buy, or the Control Plane. Anything that cannot be assigned to one of the first three has to be justified, because every row that lands in the fourth column is custom software somebody has to maintain.
| Phase 1 touchpoint | Venue | How it actually appears |
|---|---|---|
| The rep's day | ||
| Rep-specific priority list | Nutshell | A saved filtered list: owner = me, ai_tier = A, ai_next_touch ≤ today, sorted by ai_score. No new screen — the AI writes the fields, Nutshell does the sorting. |
| Next Best Action per record | Nutshell | Two custom fields on the record: the instruction, and the one-line reason behind it. |
| Pre-call brief | Nutshell Channel | Written into the record as a note when a meeting is booked; also pushed to the rep by email or Slack 30 minutes before the call. |
| Per-record AI on/off | Nutshell | A custom field the rep sets: Full · Draft only · Research only · Off. The gate reads it before every action. This is the "don't touch this one, it's mine" control you described. |
| Automatic tasks & promise detection | Nutshell | Native Nutshell tasks, created by the gate with a source tag so an AI-created task is distinguishable from a human one. |
| Live human takeover of a chat | Bought | Whatever chat product the widget runs on handles operator takeover. We supply the transcript and account context into it; we do not build a chat console. |
| Management | ||
| Daily money report | Channel | A 7am email. Top 10–25 dormant opportunities, score, reason, recommended action, each deep-linking into the Nutshell record. Same data, no dashboard to remember to open. |
| Hot-lead alerts | Channel | Slack (or SMS) with score, company, problem, recommended action and a link. SLA clock starts on delivery. |
| Neglected-lead escalation & SLA reassignment | Channel Nutshell | Escalation message to the manager; reassignment written to the Nutshell owner field. |
| Revenue-leakage detector | Channel Control Plane | Digest email of gaps — proposal sent with no follow-up, verbal yes with no contract, contract with no payment. The detector's rules and thresholds live in the Control Plane because they are policy, not CRM data. |
| End-to-end revenue measurement | Control Plane | Visitor → conversation → qualified → meeting → proposal → client → MRR, with AI-created / AI-influenced / AI-touched separated and a holdout cohort. Nutshell cannot express this; it is the attribution model. |
| Approvals | ||
| Zero-Entry CRM "Approve All" | Channel Control Plane | The one genuinely contested row. Default: a batched email with signed approve / edit / reject links, so a rep can clear a day's extractions from their phone. Control Plane holds the same queue as a fallback and for anything needing an edit. See the note below. |
| Staged reactivation batch approval | Control Plane | The first 100–250 sends need a human to read a sample and release the batch. This is an operator action, not a rep action, and it needs to show deliverability state, suppression counts and the holdout split alongside the drafts. |
| Escalation (seven triggers) | Channel Nutshell | Slack message naming the trigger, sent to whoever that trigger maps to, plus the opportunity flipped to an Escalated state with the trigger recorded on the record. Who takes it from there is handled in Nutshell, the way your team already hands work over. |
| Undoing something the system did | Control Plane | Restoring a prior value across possibly many records is not something a CRM UI is built for, and Nutshell does not retain prior values at all. The activity log and the restore live here. What is and is not reversible is set out in section 07. |
| Post-close sampling / audit review | Control Plane | A rolling sample of autonomously closed deals with the full decision trail. This is the audit layer the Canonical requires; it has nowhere else to live. |
| Transaction | ||
| Proposal delivery | Bought | Generated from your template and sent through the e-sign or proposal product, so the prospect gets a document from a system they recognise. |
| Contract signature | Bought | E-sign provider's own signing experience. We never build a signing page. |
| Deposit / payment | Bought | Processor-hosted checkout link. Card handling, retries, receipts and disputes are the processor's problem, by design. |
| Transaction state & failure recovery | Control Plane | Three external systems each know one third of the truth. Something has to join them and own the failure cases — signed but unpaid, abandoned checkout, refund request. That join is the Control Plane's transaction console. |
| Pricing guardrails & negotiation limits | Control Plane | Discount floors, approved payment terms, scope boundaries, deal-size ceiling, approved-precedent table. Versioned, effective-dated, human-edited. This is policy data and it does not belong in a CRM. |
| Kill switch & autonomy switchboard | Control Plane | Global stop, per-engine stop, per-capability autonomy level, per-segment pause. Per-record control stays in Nutshell where the rep is. |
Section 20 makes one-click batch approval a non-negotiable Phase 1 item, and its own v3 amendment says "no custom UI extension — use approved API/native Nutshell writes." Those cannot both hold unless Nutshell supports a bulk-edit-from-list-view flow we can drive, which I have not been able to confirm without the live account. There are three honest answers and Discovery picks one: an email approval digest with signed links (my default — zero new UI, works on a phone, and it is how the rep already handles their day); Nutshell bulk edit on a staging field, if the account supports it; or one screen in the Control Plane. I would rather flag this now than ship you a mock of a screen your own document forbids.
What a rep sees on Monday morning
The Canonical's rule that Nutshell stays the single system of record is the right call, and it has a concrete design consequence: everything the AI knows about a record has to be expressible as Nutshell fields. If it cannot be written to a field, a task, a note or a stage, a rep will never see it, and the system becomes a second place to look — which is exactly the failure mode you are trying to avoid.
So the first real design artefact of this project is not a screen. It is a field schema.
Sketch. Field names are placeholders; the real schema is agreed in Discovery against your live Nutshell custom-field capabilities.
The field schema, and why each one exists
| Field | Type | Written by | What depends on it |
|---|---|---|---|
ai_score | Number 0–100 | Scoring pipeline | Tiering, priority lists, hot-lead threshold, routing, SLA class |
ai_tier | A / B / C / Suppress | Deterministic thresholds | Which outreach cadence a record is eligible for |
ai_next_action | Enum, 7 values | NBA agent → gate | Rep priority list, daily money report |
ai_next_action_reason | Text | NBA agent | The Canonical's explainability rule — every recommendation states its evidence |
ai_next_touch | Date | Cadence engine | Frequency caps, list filters, contract-expiry timing |
ai_autonomy | Full / Draft / Research / Off | Human | Read by the gate before every action. Human always wins. |
ai_state | Idle / Queued / Awaiting reply / Escalated / Suppressed | Gate | Prevents two engines working the same record; drives the escalation view |
ai_escalation_trigger | Enum, the 7 triggers | Gate | An escalation names the trigger that caused it, rather than arriving unexplained |
ai_lost_reason_norm | Enum | Cleanup pass | Lost-deal resurrection, deal autopsy, scoring model |
ai_contact_status | Verified / Unverified / Departed / Bounced | Verification service | Suppression — no send without a verified path |
ai_consent_basis | Enum | Audit pass | Whether this record is contactable at all, per jurisdiction |
ai_run_id | Text | Gate | Traceability — any AI-written value on the record opens its own activity log entry |
The approved list of permissions by field
A field schema is only half the answer. The other half is a list DMA gives me: for each field, what this system is permitted to do to it. "The AI writes to Nutshell" is far too coarse a statement to build against, and I would rather be handed the boundary than infer it. It is one short document, it is yours rather than mine, and once it is signed off the gate simply enforces it. Each field carries three properties:
- May the system write it at all — deal value, close date and owner are candidates for people-only, and saying so up front is cheaper than discovering it after a bad write.
- At what autonomy — write freely, write only with approval, or propose and never write.
- What happens on collision — whether a human value is protected, and for how long.
That last property is the one that matters most in practice, and it is the mechanism behind the human-priority rule in section 07. When a person edits a field the system also writes, that field becomes theirs. The system stops writing it, keeps computing its own value, and shows the two side by side rather than overwriting. It is the honest behaviour: a rep who corrects the AI once should not have to correct it again on Thursday.
The obvious version of this leaves the lock in place forever, and over a 30,000-record database that quietly strangles the system — a year on, a third of your records are frozen by corrections nobody remembers making. So the lock has a life: it holds until the account's next substantive human touch, or a set number of days, whichever comes first. When it lapses, the system does not resume writing. It proposes, with the human value shown next to its own and the date the lock was set. The person keeps the last word; they just do not have to keep saying it forever.
Every write lands as a distinct DMA AI user rather than under a rep's credentials. That buys three things: a rep looking at a record can tell at a glance what a person did and what the system did; the Control Plane's activity log can be reconciled against Nutshell's own timeline; and the echo-suppression in section 07 has something reliable to key on.
Everything on this page assumes the live Nutshell account supports: custom fields of these types on the objects we need; API writes to them at the volume of a 30,000-record backfill plus continuous write-back; a change-notification path (webhooks, or polling we can live with); and a way to distinguish our own writes from human ones so we do not trigger ourselves in a loop. I have not verified any of this against your account. A read-only credential and two days settles it, and it is the single highest-severity unknown in the project — a polling-only or field-restricted API adds a state-reconciliation layer that no estimate from any candidate currently includes.
The surfaces you already check
Three of the Canonical's Phase 1 "non-negotiables" — the daily money report, hot-lead alerts and one-click approval — read like dashboard requirements. They are not. A dashboard is a place you have to remember to go; the whole point of the daily money report is that leadership sees it whether or not they were thinking about it. These belong in the inbox and in Slack, and building them there is both faster and more likely to actually get used.
+ 17 more · yesterday: 4 replies, 2 meetings booked, 1 escalation
Illustrative records. Every row deep-links to the Nutshell opportunity — the report is a pointer, never a second system.
The escalation names its trigger. "Add as precedent" is the mechanism that makes the system get less needy over time — see section 09.
Where an escalation is sent
The Canonical routes all seven triggers to Josh, and reading it closely that is a heavier load than it first appears: Josh also takes every trust-building video call, by a standing rule set outside the seven triggers. One person as the sole destination for every exception is a bottleneck, and the failure is silent — the escalation fires, that person is mid-call, and the deal sits.
The fix is small and it is a mapping, not a process. Each of the seven triggers points at one or more destinations, and that mapping is a setting in the Control Plane that DMA edits directly — no deploy, no ticket to me. That is the entire extent of what the system needs to know.
Deliberately, it knows nothing more. It does not track who is on holiday, run a timer, chase an acknowledgement or re-route to a backup, because Nutshell and Slack already do the human half of this better than a rule I would write in advance: the opportunity is flipped to Escalated and assigned, and from there people tag each other, watch a colleague's queue while they are out, and hand things on. Encoding coverage rules into this system would freeze an arrangement your team changes every week, and it would quietly become another place to keep up to date.
What the system owes you instead is visibility, which costs nothing to provide: escalated opportunities are a state in Nutshell, so "what is escalated and how long has it been sitting there" is a saved list rather than a feature. If DMA later wants a clock on that, it is a Nutshell automation over a field that already exists — not something to build here now.
One route does need to be populated at all times, and it is the only hard requirement: the prospect who asks for a human. That trigger has no confidence threshold and no deferral, so its mapping cannot be empty. Everything else can be left to the team.
If the default approval surface is a link in an email, that link is a credential and has to be treated as one. Each is single-use, short-lived, bound to one named approver, and bound to the exact version of the thing they were shown — so if a draft changes between the digest going out and the click landing, the link no longer approves anything and the item returns to the queue. Approving is the only thing these links can do; they are not a way into the Control Plane. A forwarded email should never be able to release a batch of outreach.
Real booking is a Phase 1 item and it is a solved problem: the agent reads availability and writes the event through the calendar API, against the routing rules for service lane and deal size. The one design decision worth making early is that the agent books into a real calendar, not a booking-link handoff — the Canonical is right that a booking link is a drop-off point, and once the agent has the prospect's attention it should not let go of it.
The services, and who should own the accounts
Every one of these is a category where building is strictly worse than buying. My recommendation across the board is that DMA holds every account directly — you own the relationships, you see the spend, and if you ever change development partners nothing has to be untangled. I would quote integration effort only, and exclude all pass-through fees and model spend from any price. That also keeps your comparison between candidates honest, since a bundled number hides which is which.
| Need | Shape of the answer | Integration effort | Notes |
|---|---|---|---|
| E-signature | PandaDoc / DocuSign class | Medium | Needs template + merge fields + completion webhook. Choice partly depends on whether you want proposal and contract in one product. |
| Payments | Stripe class | Medium | Hosted checkout, not custom. Retry, dunning and dispute handling come with the product and should not be reimplemented. |
| ESP + sending domain | Dedicated sending domain, SPF/DKIM/DMARC, warm-up | Medium | Must be configured and warmed before any reactivation send. This is on the critical path for the first revenue milestone. |
| Email verification | Verification + finder service | Low | Gates every send. Directly determines how much of the 30K is reachable at all. |
| Firmographic enrichment | Clearbit class | Low–Medium | Feeds scoring and personalisation. Per-record cost across 30K is a real budget line. |
| SEO / site analysis | Likely covered by DMA's existing tooling | Low | Feeds the Mini-Consultant. Worth confirming what you already pay for before buying anything new. |
| Website chat widget | Extend existing, or embed new | Unknown | Depends entirely on your website platform, which I do not yet know. Also determines how live human takeover works. |
| Calendar | Google / Microsoft API | Low | Direct API, not a booking-link product. |
| Alerting | Slack, optionally SMS for hot leads | Low | — |
| Model provider | Provider-flexible; routing by task | — | Cheap models for bulk extraction and classification, strong models for customer-facing copy and negotiation. Costed below. |
What the model spend actually looks like
These are estimates, and I want to be exact about what kind. The unit prices are published list prices. The token volumes per task are my engineering estimates for work of this shape, not measurements — I have not run your data through anything. The monthly volumes are assumptions about DMA, and you should correct them. Everything below can be recalculated once any of those three change, which is the point of showing the arithmetic rather than a single number.
Two model tiers, chosen per task. A cheap, fast model does the bulk reading — 30,000 records, transcripts, inbound replies, classification. A stronger model writes anything a prospect will read. Mixing them is not a cost trick; sending a reactivation email through the cheap model to save a third of a cent is a bad trade against the reply rate.
| Unit of work | Tier | Estimated cost each |
|---|---|---|
| Read and score one dormant record | Bulk | $0.004 – $0.020 |
| Write the "why now" line that travels with a reactivation candidate | Writing | $0.008 – $0.014 |
| Draft one personalised reactivation email | Writing | $0.010 – $0.025 |
| Classify an inbound reply and draft the answer | Both | $0.012 – $0.020 |
| One whole website conversation | Writing | $0.060 – $0.150 |
| One live site review for the Mini-Consultant | Writing | $0.030 – $0.060 |
| One call transcript into CRM fields, tasks and a follow-up draft | Both | $0.020 – $0.035 |
| Re-score a record after it changes in Nutshell | Bulk | $0.003 – $0.008 |
| Operating tempo | Assumed volume | Estimated model spend |
|---|---|---|
| Reactivation at a low cadence | 800 drafts · 100 conversations · 60 site reviews · 50 calls · 1,000 re-scores | $40 – $90 |
| Reactivation at full cadence | 2,000 drafts · 300 conversations · 200 site reviews · 150 calls · 3,000 re-scores | $90 – $200 |
| Full cadence over a heavy inbound month | 5,000 drafts · 1,000 conversations · 500 site reviews · 300 calls · 6,000 re-scores | $230 – $520 |
| One-time | Basis | Estimate |
|---|---|---|
| First full pass over 30,000 records | 30,000 × the per-record range, plus a written reason for each record that clears the tier B threshold. How many that is, is one of the things the first pass tells you — the estimate spans a quarter to a half of the database | $200 – $750 |
| The same pass, submitted as batch work | Non-urgent bulk work runs at half price, and the audit is the definition of non-urgent | $110 – $400 |
| Prompt tuning and regression runs before go-live | Repeated evaluation passes over a fixed set of DMA examples | $50 – $250 |
Model spend for a system of this shape lands in the low hundreds of dollars a month. The data vendors in the table above do not. Firmographic enrichment across 30,000 records is a one-time charge in the hundreds to low thousands before any ongoing refresh; email verification adds a similar one-time hit; an ESP with a dedicated domain, an e-sign seat and SEO tooling are each a monthly line in the tens to hundreds; payment processing is a percentage of revenue. Realistically the third-party stack costs several times what the AI does, and it is the line that actually needs a decision from DMA. A quoted running cost that counts only model tokens is measuring the smallest item on the bill.
None of these depend on a model behaving. A hard monthly ceiling that refuses further calls and falls the website chat back to a plain capture form. A separate daily ceiling per subsystem, so a runaway bulk job cannot starve the live chat. A per-conversation cap that hands to a human rather than spending indefinitely on one visitor. Rate limits per visitor and per conversation length. And the expensive reads are keyed on their inputs rather than on the clock — the same content fingerprint the adapter uses in section 07 decides whether anything needs looking at twice, so the bill tracks how much your data actually moved rather than how often the system woke up. That is the whole reason the recurring number is a fraction of the one-time one: after the first pass, you are paying for change, not for volume.
The only thing here that has to be purpose-built
Everything in the previous three sections was Nutshell, a channel, or a product you buy. What is left is the Control Plane, and the test I applied to every item in it is simple: could this live in Nutshell, in an inbox, or in a product we can buy? If yes, it is not in here. What survives that test is eight things, and none of them are sales work.
This matters for the Canonical's "no custom approval UI" rule. The rule is about not making reps learn a second CRM, and this design honours it — a rep can do their entire job without ever opening the Control Plane. What the rule cannot mean is that a system with autonomous pricing, contracts and payments has no operator console at all, because then nobody can turn it off, nobody can change a discount floor, and nobody can answer "why did it do that?" six weeks later.
Action queue — the operator's screen
Sketch of placement and content. Note what is not here: no lead list, no pipeline view, no contact records. Those are Nutshell's, and duplicating them is how you end up with two systems of record.
What happens when everything moves at once
This is the hardest part of the project, and the part most likely to be quietly missing from a system that demos well. An agent is drafting a reply, a rep is editing the same record, a prospect's email arrives and Nutshell fires a webhook — all inside the same few seconds. Who wins, what gets thrown away, and what must never be allowed to happen twice? Here is the mechanism, in full.
No agent ever writes directly. An agent produces an intent: a proposed action, the snapshot of state it was computed from, the evidence behind it, and an idempotency key. The intent goes on a queue. The gate is the only code in the system that touches Nutshell, sends mail, or calls the payment and e-sign providers, and it runs these steps in order every time.
ai_autonomy field, suppression list, unsubscribe, frequency cap, quiet hours, active-client check, open support or billing issue. Circuit breakers sit here too: bounce and complaint rates crossing their thresholds pause sending automatically rather than waiting for someone to notice on a dashboard, because by the time a deliverability problem is visible in a weekly review the sending domain is already damaged.
→ Blocked: log the reason, leave the record untouched.
Escalated.
ai_run_id stamped on the Nutshell record, so any field a rep is looking at can be traced back to the reasoning that produced it. Reversible classes of action then sit in a short hold window before they take effect, which is the only free chance anyone gets to stop them.
Writing exactly once, against an API that may not help
The textbook answer is an idempotency key: attach a unique token to the write, and the far end guarantees it applies once no matter how many times you send it. If Nutshell supports that, we use it and this problem is over.
It may not, and most CRM APIs do not. The failure that matters is narrow and specific: we send a write, the connection dies before the response arrives, and we genuinely do not know whether it landed. Retrying blindly gives your prospect two identical notes, or two tasks, or a duplicated activity. So the adapter carries its own answer. Every created object gets a deterministic fingerprint derived from its content and the run that produced it, written into the object itself. Before any retry, the adapter looks for that fingerprint. Found means the first attempt succeeded and the retry is recorded as a no-op. Not found means it genuinely failed and the write proceeds. For field updates the problem is easier — if the field already holds the value we meant to write, we are done.
I am flagging this rather than assuming it away because it is exactly the kind of thing that works perfectly against a simulator and then does not survive the real API.
When Nutshell tells us something changed
Change notifications are treated as signals, never as data. A webhook tells us a record is worth looking at; it never tells us what the record now says. The system goes and reads the live record before doing anything with it.
That one rule disposes of a whole family of problems at once. Events arriving out of order stop mattering, because whichever arrives last still triggers a fresh read of the same current truth. A replayed or duplicated event costs one redundant read and is otherwise inert, and duplicates are dropped on event identity anyway. And a notification whose payload is stale — which is normal under load — can never roll a record backwards, because we never trusted the payload.
Our own writes come back as notifications too. The activity log knows what the system just wrote, so an echo of our own change is recognised and consumed silently rather than waking an agent. Without that, the system responds to itself, and a small loop becomes a large one very quickly.
The nightly argument with Nutshell
Everything above assumes we are told when things change. Assume we are not. Notifications get dropped, a polling window gets missed, someone does a bulk import, an integration nobody mentioned writes to the same fields. Any working mirror of a system you do not control needs a reconciliation loop, and it needs to do something when it finds a problem.
So a full pass compares the mirror against Nutshell nightly, and a continuously running sample checks a small slice throughout the day — because a bad deploy at 9am should not have until midnight to do damage. Discrepancies are sorted into two kinds, which matter differently. A value we wrote that no longer matches means a person changed it and we missed the notification: the mirror is corrected and that field becomes theirs, exactly as if we had seen the edit. A value neither side can account for is a genuine integrity problem and is the one that should make somebody nervous.
Drift is measured per field class rather than per record, because the thresholds are not the same: a percent of contact-detail churn overnight is ordinary, while the same drift in deal stages or values is not. When a class passes its threshold, writes for that class pause automatically and the queue holds rather than drains. Nothing is dropped, nothing is lost, and someone has to look before it resumes. The alternative — a system that keeps confidently writing into data it has quietly lost track of — is how a CRM integration turns into a data-recovery exercise.
What can actually be undone
"Undo" is the first thing anyone asks for and the word covers two very different things, so it is worth separating them plainly.
Writes into Nutshell can be put back — but not by Nutshell, and I checked that rather than assuming it. Nutshell's own Audit Log is an Enterprise-plan admin screen covering logins, bulk edits and list exports; it cannot be exported and has no API behind it. Its change feed is more promising but stops short of what an undo needs — Nutshell's own guidance is to treat a change notification as a prompt that an entity moved, not as a description of what moved, and the sample payload ships an empty changes array with only the entity's current state. Nothing in Nutshell can tell you what a field held before it was overwritten. So prior values have to be retained on our side, and the only real design choice is how much apparatus that justifies.
My answer is: almost none, because the Canonical already requires the expensive part. Every AI decision has to be traceable and costed, which means the Control Plane keeps an activity log regardless — one entry per action, with the agent, the rule that allowed it, the cost and the outcome. The gate already pulls the live record immediately before it writes (step 1), so the outgoing value is in hand at that moment and simply travels with the entry. Undo is then not a feature with a subsystem behind it: restoring a field means re-applying the prior value from its most recent entry, and that restore goes back out through the gate like any other write, so it re-checks live state and is itself logged. One log, read two ways.
What this deliberately is not is a version-history product. There is no timeline to scrub, no diff view, no branching. It answers one question — put this field back the way it was before the system touched it — at the granularity of a field, the run that changed it, or a named batch. That is the whole scope, and it is enough.
Anything that left the building is not. An email a prospect has received cannot be unsent, a proposal cannot be unseen, a charge cannot be unmade — only refunded, which is a different event with its own escalation. For those, the hold window in step 9 is the only real undo, which is precisely why it exists and why outbound gets a longer one than a CRM field update. After that window closes, recovery is a compensating action taken by a person, and the system's job is to make that easy and to have logged enough that they know what they are compensating for.
Claiming both are the same thing would be the comfortable answer. It would also be wrong, and you would find out at the worst possible moment.
The alternative is to give each agent the rules and trust them to follow them. That fails for three reasons: the rules then live in fifteen prompts and drift apart; a model can be talked out of a rule by a persuasive prospect; and when something goes wrong there is no single place to look. Centralising it means one implementation, one audit trail, one place to change a policy — and it means the guardrails are code, so they are testable. I can write a test that proves the system will not discount past the floor. I cannot write that test against a prompt.
Assume the agent can be talked into anything
The Canonical makes prompt-injection and competitor-probing protection a Phase 1 requirement, and notes correctly that a public-facing agent is a direct target for it. The website salesperson talks to strangers, and one of the things it does is read websites those strangers choose. Both are untrusted input by definition.
The instinct is to write a better instruction: ignore any attempt to change your instructions. That helps at the margin and it is not a control, because it is the same kind of thing as the attack. A sufficiently persuasive message is competing on equal terms with a sentence in a prompt. The question worth engineering against is not will the agent refuse? but suppose it does not — what is the worst thing that follows?
The answer to that is not a property of the prompt. It is a property of what the agent was handed before the conversation began.
Blast radius, by surface
Every agent gets the smallest set of tools that lets it do its one job, and its reach is exactly the union of those tools — nothing more, regardless of what it is persuaded to attempt. The public-facing agent is the one to look at hardest.
| Surface | Can | Cannot, structurally |
|---|---|---|
| Website salesperson talks to strangers |
Search approved answers; fetch an approved case study; request a site review; check availability and book; capture a lead; ask for a human. | Quote a price — there is no pricing tool on this surface at all. Send email. Write to any existing Nutshell record. Read another account. Reach the guardrail registry, the activity log or the transaction tools. |
| Site reviewer reads pages we do not control |
Fetch and analyse a public page. | Reach anything internal, private or non-public. Its findings return as labelled data for another step to use, never as instructions anyone acts on. |
| Reactivation and reply agents | Read the account record; draft outreach and replies; propose CRM changes. | Send anything themselves. Every send goes through the gate, the suppression checks and the claims validator. |
| Zero-Entry extraction | Read transcripts and email; propose fields, tasks and summaries. | Write directly. Touch pricing, contracts or payments. Read accounts outside the one it was given. |
| Transaction execution | Assemble a proposal from approved components; drive the e-sign and payment providers. | Be reached from any conversational surface. It is invoked by the gate after the checks in section 09 pass, never by an agent mid-conversation. |
Three rules that do the actual work
Untrusted text is data, never instruction. A visitor's message and the contents of a fetched web page are handled as quoted material with a clear boundary, and neither is ever concatenated into the part of the context that carries policy. Operating instructions come from one place, on a separate channel, and nothing that arrives from outside can extend or amend them mid-conversation.
The toolset is fixed before the conversation starts. An agent cannot acquire a capability during a conversation, because tool grants are configuration attached to a specific agent version, not something negotiated at runtime. There is no sequence of messages that gets the website chat an email tool.
Least privilege goes all the way down. Each capability group holds its own credentials against the database and against Nutshell, scoped to what it actually needs. A flaw in the component that handles public chat does not become read access to your 30,000 records, because that component was never able to read them.
Containment limits what the agent can do. Competitor probing is about what it can be induced to reveal — pricing logic, the playbook, how leads are scored, what it was told about a competitor. Two things hold there. The scoring model, the guardrail registry and the internal playbook are not in the public agent's context to begin with, so there is nothing to extract; it can reach approved public content and nothing else. And the claims validator in step 7 of the gate runs on the way out as well as the way in: anything the agent produces that asserts a figure, a ranking, a result or a guarantee has to resolve to approved source material, which catches both the invented claim and the leaked internal one. Output validation is the control that does not care how clever the input was.
Which decisions a model is allowed to make
The sharpest form of this question is how "no approved precedent → Josh" gets implemented without simply asking a language model whether something feels unprecedented. That one is answered on its own below.
| Decision | Decided by | Mechanism |
|---|---|---|
| Whether a record may be contacted at all | Code | Consent basis, suppression list, unsubscribe, frequency cap, quiet hours, bounce state, active-client check. No model input. |
| Tier boundaries and eligibility | Code | Fixed thresholds over the score. Changing a threshold is a versioned config change with an author. |
| The opportunity score | Hybrid | The model extracts features from free-text notes, emails and transcripts; the score is a formula over those features with versioned, human-approved weights. The number is reproducible; the reading of the prose is not, which is exactly the right split. |
| The price itself | Code | The model never produces a number. It selects services and quantities; code prices them from the approved rate card. See below. |
| Discount | Code | A lookup against the approved ladder. The model may select a rung it can justify; it cannot invent one, and anything past the floor escalates rather than being clamped to it. |
| Payment terms, contract template, term length | Code | Allow-list. Anything not on it is by definition non-standard, which is escalation trigger 2. |
| Deal-size ceiling | Code | Numeric comparison. |
| Whether to send a proposal | Code | A gate over a model-drafted artefact. All seven triggers plus the claims check, all evaluated in code. |
| Whether there is approved precedent | Code | A table lookup. See below — this is the one worth spelling out. |
| Refund or cancellation | Code | Always escalates. No exceptions, no confidence threshold. |
| Payment retry and dunning | Code | The payment processor's own logic, driven by a state machine. Explicitly not an agent — an agent has no business deciding when to re-charge a card. |
| Which sales methodology to run | Model, bounded | Selection from a fixed set of five, driven by signals the system already produces. The selection is logged so it can be reviewed against outcomes. |
| Tone, phrasing, objection responses | Model | Bounded by the playbook, the objection library, and the claims validator on output. |
| Summarising, extracting, classifying | Model | Structured output against a schema. A response that fails schema validation is a failed run, not a partial write. |
| Merging duplicate records | Human | Detection is automated, merging is not. Merges are destructive and hard to unwind — your own Canonical corrected this one, and it is right. |
The system does not compose prices
There is a weak version of a pricing guardrail and a strong one, and the difference is worth being explicit about because they sound alike.
The weak version lets the model write a number and then checks whether that number is allowed. It mostly works. It also means a figure exists, inside the system, that no rule produced — and the entire safety of the arrangement now rests on the validator being complete. Every gap in it is a price nobody approved.
The strong version removes the model from the job. It decides what the prospect needs: these services, this scope, this term. Code turns that into money, line by line, from the rate card, applying the discount ladder. The number is computed, reproducible, and attributable to a specific version of a document DMA wrote.
The engineering that enforces this is unglamorous and effective: the model's output schema has no price field. It cannot emit what it has no slot to emit. And on the public website surface there is no pricing capability at all — an inbound visitor asking "what would this cost?" gets a range only if approved public ranges exist for that service, delivered by code, or gets a conversation about scope and a booked call. Nobody has to trust the agent to be careful with your pricing, because it was never handed it.
"No approved precedent" as a lookup, not a judgement
Asking a model "does this seem unprecedented?" is unanswerable, because the model has no reliable memory of what DMA has approved before and every incentive to be agreeable. So the system does not ask it.
Instead, every autonomous commercial action has to resolve to a precedent record. A deterministic classifier reduces the action to a key made of a fixed set of attributes — service mix, deal-size band, contract term, payment terms, contract template version, discount rung, and any non-standard scope flags. The model's only job is to fill those slots from the conversation; it does not get a vote on what the slots mean. The gate then looks the key up in the approved-precedent table, which is seeded in Discovery from your historical closed-won deals and your written pricing policy.
- Key found — this shape of deal has been approved before. Proceed autonomously.
- Key not found — escalate to Josh, naming the attribute that had no precedent. Nothing is sent.
- Josh approves — the key is added to the table with the approver's name and the date on it. The next deal of that shape proceeds without an escalation.
That last line is the important one. The system starts conservative and becomes autonomous as a direct function of decisions a human actually made, rather than as a function of how confident a model happens to sound. And because the table is data, you can look at it, argue with it, and take entries out.
Who is allowed to change the rules
All of the above is only as good as the change control around it. A discount floor that any one person can quietly edit is a suggestion. So every load-bearing setting — the guardrail registry, the precedent table, the escalation routing, the field permissions, and the agents' own instructions — is treated as code rather than as configuration: an edit is a proposal that sits inert until a second named person signs it off, and it carries that person's name for as long as it is in force. Authorship and authorisation are never the same signature, mine included.
Two properties follow that matter more than the rule itself. Settings are effective-dated rather than merely current, so a decision taken in March is replayed against the floors and terms that were in force in March — without that, your audit trail explains what happened using rules that did not yet exist. And because each signed-off state is retained, withdrawing a bad change is a one-line operation with a name attached, instead of a conversation about what the number used to be.
This is also the Canonical's human-reviewed model changes requirement, which asks that the system never silently rewrites its own sales strategy. Made concrete, that is exactly this: a change to the strategy is a reviewed commit with an author, an approver and a date.
The Control Plane says the AI hands a deal to Josh when one of seven conditions is met. Trigger 7 is "no approved precedent." Trigger 1 is "deal exceeds the largest previously-approved deal size" — and then says "Not blocked — the first deal at a new tier proceeds and gets sampled review after close." A deal above every previously approved size has no approved precedent by definition, so trigger 7 says stop and trigger 1 says go, on the same deal. This is not a nitpick: it decides whether your single largest-ever deal gets sent autonomously. My recommendation is that trigger 1 holds and trigger 7 yields — a new size tier is exactly the moment a human should look — but that is DMA's commercial call, not mine, and it needs making before the transaction layer ships.
Conversation to collected money, including when it breaks
The four v3 additions — methodology orchestration, autonomous proposal sending, e-sign execution and deposit capture — are the newest and least specified part of the Canonical, and they are where the real risk sits. The happy path is straightforward. The value is in the branches.
Four rules specific to money
The gate in section 07 governs everything. These four apply only to the transaction path, because the cost of getting them wrong is different in kind.
- What was approved is what goes out. The approved document is fingerprinted at the moment of approval and checked again at the moment of sending. If anything about it changed in between — a regenerated draft, a rate card edited underneath it — the send stops and returns for approval. An approval is for a specific artefact, not a general blessing.
- Whoever requested it does not release it. On anything that binds DMA or moves money, the person who asked and the person who approved are two different people. This is the same separation as the rule changes in section 09, applied where it matters most.
- The provider tells us, and we check anyway. Signature and payment callbacks are verified as genuinely from the provider before anything acts on them, and a reconciliation job asks each provider for the current state of anything that has been open too long. A callback that never arrives is a normal event, not an exception — designing as though it cannot happen is how a paid deal sits in Negotiation for a fortnight.
- Card data never touches this system. Payment runs through the processor's own hosted checkout. This is not primarily a cost decision; it keeps DMA's compliance surface where it belongs.
The branches that actually matter
| What goes wrong | What the system does | Where a human sees it |
|---|---|---|
| Signed, payment fails | Opportunity moves to a distinct Signed — unpaid state so it never sits misleadingly in Closed Won. The processor runs its own retry and alternative-method flow. The system sends a short, non-accusatory follow-up and tries an alternate method. |
Transaction console; Slack to the owner. Trigger 3 escalates on any dispute. |
| Signed, checkout abandoned entirely | Timed follow-up sequence with a fresh link, then hard escalation. It does not time out and quietly drop the deal — you have a signed contract and no money, which is a human problem. | Slack escalation; leakage digest; console. |
| Prospect asks for a human | Immediate handoff, full context, no further autonomous messages on that opportunity. Trigger 4, no confidence threshold, no exceptions. | Slack to Josh; opportunity flagged. |
| Video call scheduled | Standing rule from Jeev, independent of the triggers: Josh takes the call. The agent prepares the prospect, briefs Josh, and resumes afterwards from the transcript. | Calendar; pre-call brief. |
| Refund or cancellation requested | Always escalates. The system takes no position and sends nothing. | Slack; console. |
| Rep intervenes mid-flow | Human priority. Pending intents on that opportunity are discarded and re-planned from the state the rep left behind. | Nothing to see — which is the point. |
Autonomous proposal sending is Phase 1. Its listed dependency, proposal data pre-fill, is Phase 2, as is proposal first draft — which the document also notes is a duplicate of Section 11's automatic proposal draft, itself Phase 2. So the capability that sends proposals ships a phase ahead of the capabilities that build them. Separately, e-sign execution is Phase 1 and depends on "approved contract templates," but no capability in the document produces those; they are DMA authoring work with no named owner. Both are resolvable — pull the pre-fill and draft items into Phase 1, and name a template owner — but they need resolving before this cluster can be scheduled, let alone priced.
What is demonstrably working, and when
A milestone view at Day 10, 20, 30 and 45, with the precondition for each stated plainly, because a milestone whose precondition has not been met is a date nobody should trust. Days are counted from build day one; Discovery precedes day one and runs five to eight days of its own.
The sequencing principle: the 30,000 records are the fastest path to real revenue, so they come first, and nothing downstream gets turned on until the thing above it is trusted. We are not turning the whole pipeline on at once and hoping.
| By | Demonstrably working | Requires |
|---|---|---|
| Day 10 Foundation |
Nutshell adapter proven against the live account — reads, writes, rate limiting, retry, exactly-once handling, change detection, echo suppression. All 30K mirrored, incremental sync running, and the nightly reconciliation with its drift thresholds already switched on. Activity log with prior-value capture, restore, kill switch and autonomy switchboard live from the first day, not retrofitted. First real data-quality numbers on the 30K: record age, field completeness, lost-reason population, contact validity, duplicate density.
Shown as: a live sync you can watch, a write made and then put back from the activity log, and the first honest numbers behind the reactivation thesis. |
Nutshell API credential. Agreed field schema. Nothing from DMA beyond access. |
| Day 20 Database becomes a work queue |
Scoring pipeline over all 30,000: score, tier, Next Best Action and its reason, written back to Nutshell under the approved list of permissions by field. Rep priority list as a saved Nutshell list. Daily money report in production. Hygiene passes run: duplicate detection, dead contacts, lost-reason normalisation, stale deals, orphaned opportunities. The scoring model backtested against deals DMA has already won and lost, so we can say whether it would have ranked your real wins near the top before we ask anyone to trust it on a dormant record.
Shown as: Josh opens Nutshell and has a ranked list with reasons attached; leadership gets the 7am email; and the backtest either supports the model or tells us what to fix. |
Day 10. Two or three working sessions on what "good lead" means, because the first scoring pass will disagree with you and that disagreement is the useful part. |
| Day 30 First real revenue motion |
Sending domain warmed and authenticated. Verification and suppression enforced. Personalised reactivation drafts released in a staged first batch of 100–250 with human release. Replies ingested, classified, answered, objections handled, meetings booked — all through the Commit Gate. Hot-lead alerts and the SLA clock live. Holdout cohort in place so lift is measurable rather than assumed.
Shown as: real sends to real people, real replies handled, first meetings on Josh's calendar, measured against a control group. |
Day 20. Sending domain. Consent position settled. ESP and verification accounts. A decision on what a "successful" reactivation gate means — a send, a reply rate, a booked meeting, or closed revenue. |
| Day 45 Inbound engine + transaction path |
Website agent live: conversation, Mini-Consultant, objection handling, live re-scoring, booking, hot-lead alert, Nutshell write-back. Zero-Entry CRM's five items running off existing Zoom transcripts. Transaction path complete end to end in staging against an e-sign sandbox and payment test mode — proposal, signature, deposit — with all seven triggers wired and the precedent table seeded, and the failure branches from section 10 deliberately exercised.
Shown as: a full transaction walked end to end on a test deal, and then walked again while we break it on purpose. |
Day 30. Website platform access. Pricing policy written. Contract and proposal templates. Approved proof and objection libraries. E-sign and payment accounts. |
Day 45 is not "autonomous closing is on"
I want to be plain about this because it is where I expect my answer to differ most from a competing bid. Having the transaction path working is not the same as trusting it with your customers and your bank account. Autonomy gets raised as a deliberate, reversible ramp, per capability, and the ramp is measured in weeks of real deals rather than a date on a plan.
Building this is the predictable half. The unpredictable half is the iteration: the first scoring pass will rank one of your best-ever clients a C, and the useful question then is whether the model is wrong or whether it has found something about your intake that you want to change. That conversation happens repeatedly, it needs you in the room, and it is the difference between a system that runs and a system that earns. I have built this shape of thing before — most recently against Pipedrive, which we use ourselves — and the build is not what takes the time.
What the Canonical does not yet resolve
The Canonical's own last page says it does not select vendors, design the schema, or produce a budget and timeline. It is a strong decision document and I have treated it as binding. These are the places where following it precisely produces a question rather than an answer — raised now, because finding them in week three is considerably more expensive.
| Item | The tension | Effect if unresolved |
|---|---|---|
| Scope of Phase 1 | 92 rows are marked Phase 1, of which at least four pairs are duplicates of each other. Is the release all 92, or the eleven non-negotiables plus the four v3 transaction items? | Blocks pricing Roughly a fivefold swing in effort. The largest single variable in any bid you receive. |
| Nutshell API reality | Every capability rests on two-way access. Write support for the assumed objects and fields, change notification versus polling, per-write idempotency support, bulk-edit behaviour, and rate limits under a 30K backfill plus continuous write-back are all untested. | Highest severity A polling-only or field-restricted API changes the sync design rather than adding to it. Worth saying plainly: this is the one question no prototype can answer. A system demonstrated against a stand-in CRM — ours or anyone's — has proved that the software works against the assumptions its authors made about Nutshell, which is a different claim from working against Nutshell. Two days with a read-only credential settles it, and it should be settled before anybody signs anything. |
| "No custom UI" vs. four Phase 1 surfaces | Batch Approve All, the daily money report, rep priority lists and hot-lead alerts are all Phase 1, and the document forbids a custom approval UI. Section 02 proposes an answer for each; they need confirming against the live account. | Decides whether there is a web application in Phase 1 at all. Probably the second largest line item available. |
| Triggers 1 and 7 | A new deal-size tier has no approved precedent, so one trigger says proceed and the other says escalate, on the same deal. | Decides how your largest-ever deal is handled. Needs a decision before the transaction layer ships. |
| Proposal dependency chain | Autonomous proposal sending is Phase 1; the pre-fill and draft capabilities it depends on are Phase 2. | Either pull two items forward or move the sending item back. Straightforward once someone decides. |
| Shadow pipeline references | v3 reverses the shadow pipeline to NO, but estimated latent pipeline value and promote when evidence becomes strong both still list it as their dependency. | Low severity. The amendment explains the intent; the dependency rows were not updated. |
| Transcript extraction phasing | Automatic call summary and CRM field extraction are Phase 1 non-negotiables; extract facts from call transcripts is Phase 2 and rated High risk. They read the same source and do substantially the same thing. | Either the Phase 1 items carry the Phase 2 legal gate, or the phasing needs correcting. This one is worth getting right — it is the only High-risk row in the document. |
| Rep coaching phasing | The v3 amendment says "PULLED FORWARD"; the row still reads Phase 3. | Low severity, but it is a governance-sensitive capability and should not be ambiguous. |
| Content that does not exist yet | Objection library from real sales history, case studies tagged by vertical, approved competitive positioning, approved proof points, sales playbook, contract and proposal templates. All are Phase 1 dependencies. None has a named owner. | Most likely to move dates This is DMA subject-matter time, not engineering time, and in my experience it is the most underestimated item in projects of this shape. |
| Pricing policy and guardrails | The Canonical's own dependency list requires documented pricing policy before guardrails, negotiation limits or autonomous sending can be built. The guardrails are load-bearing throughout and never specified. | The transaction cluster cannot be estimated or scheduled until this exists. It is DMA's commercial judgement; I can run the sessions that get it written down. |
| Compliance positions | Consent basis and jurisdictions for contacting 30,000 dormant records; comfort with an agent binding the company and taking payment; scope of call-recording use; a data processing agreement covering records and transcripts passing through a model provider. | A restrictive answer on outreach consent shrinks the contactable universe, which is the core value thesis. Needs an owner and a date, not a resolution from me — I do not give legal advice. |
| Acceptance criteria and baselines | The Phase 1 scorecard names seven metrics and sets no targets. There is no baseline close rate, average deal size, cycle length or inbound volume to measure lift against. | Determines how much tuning sits inside a fixed price. A revenue-based milestone gate carries materially more delivery risk than a functional one and should be priced as such. |
Dependencies, and when each one bites
These are the items I am assuming DMA provides or resolves, and which milestone moves if one of them slips. Short list, honest answer.
| Needed | By | If it slips |
|---|---|---|
| Nutshell API credential, ideally read-only first | Before day one | Everything. This gates the whole project and two days of testing settles the largest unknown in it. |
| Agreement on what a good lead looks like | Day 10–15 | Day 20 scoring. The first pass will disagree with you; those sessions are where it gets fixed. |
| The approved list of permissions by field — what this system may write, and what stays people-only | Day 10 | Nothing, if we have it. It is an hour of someone's time and it prevents the class of problem nobody forgives. |
| Escalation routing — which people or channel each of the seven triggers goes to | Day 25 | Day 30, and only briefly. It is a short list you can change yourself afterwards; the one route that cannot be empty is "prospect asked for a human". |
| A second named approver for rule and guardrail changes | Day 30 | Nothing directly, but without it the four-eyes control in section 09 is decorative. |
| Consent position on contacting dormant records | Day 20 | Day 30 outright. Nothing goes out without it. |
| Sending domain and ESP account | Day 20 | Day 30. Warm-up takes calendar time that cannot be compressed by working harder. |
| Written pricing policy, discount floors, payment terms, deal ceiling | Day 30 | The entire transaction cluster. This is the one I would start on today regardless of who you hire. |
| Approved contract and proposal templates | Day 35 | Day 45 transaction path. No capability in the Canonical produces these. |
| Objection library, proof points, case studies, competitive positioning | Rolling, from day 15 | Quality of every customer-facing message. Does not block a demo; absolutely determines whether the system converts. |
| Website platform access and chat widget decision | Day 30 | Day 45 inbound engine only. The reactivation engine is unaffected. |
| Zoom transcript access and retention scope | Day 30 | Zero-Entry CRM items at day 45. |
| E-sign and payment accounts in DMA's name | Day 35 | Day 45 staging walkthrough. |
| A named person for weekly review | From day one | Nothing on paper, everything in practice. This is worth considerably more than it sounds. |
Why I sent you this instead of a demo
A generated proof of concept answers the question "can something that looks like this exist?" The answer is yes, and it has been yes for about two years. It does not answer the questions that decide whether this project works: whether your Nutshell account supports the writes the design assumes, where a rep looks when the AI has done something, what stops two agents writing to the same opportunity, how a discount floor is enforced rather than suggested, what a clever visitor can talk the website agent into doing, how you take back something it got wrong, and what happens to a signed contract that was never paid for.
These three documents answer those. They are not complete, and I have said clearly where they are not. Treat this as an early plan rather than a settled one: I have reviewed and shaped every position in it and will defend each one, and I would still expect Discovery to revise parts of it thoroughly. If you hire me, this is the shape of what I will build. The real design will be better than this one, because it will have your live Nutshell account, your pricing policy and your sales history behind it rather than my inferences about them.
Happy to walk any section of this on a call, and equally happy to be told I have read something wrong in the Canonical — it is a dense document and I would rather be corrected now.
Prepared against the DMA AI Revenue System v3 Canonical and the finalist confirmation questionnaire.
Part Three sets out how I would build it.
Illustrative screens in this part; no live data anywhere.
What this document decides, and what it cannot
Part Two answers where does a human look, and what keeps the system honest. This part answers what gets built, in what order, with which tables and which prompts. It is the half an engineer picks up on day one, and it assumes you have read the other.
Everything here is a proposal made without access to the live Nutshell account. Where a decision depends on something I have not been able to measure, it is marked and carried to section 15 rather than guessed at silently. Schema is given as real DDL because a schema written in prose is a schema nobody has checked; treat the column names as a first draft and the shapes as the argument.
Two conventions carry through. Machine marks something the system does without a person. Human marks a point where a person is required, and each one is there because a rule put it there, not because the design was nervous.
Fifteen agents, one queue, one gate, one log. Agents never touch the outside world; they emit intents. The gate is the only code that writes to Nutshell, sends mail, or calls the e-sign and payment providers. Everything the system did is one table, read two ways — as an audit trail and as an undo. If you only take one thing from this document into a code review, make it that the gate has no bypass.
Boring where it can be, deliberate where it must
The selection rule was: minimise the number of systems that can be independently wrong at 3am. That argues for one language, one datastore and one deployment target until measurement says otherwise.
| Layer | Choice | Why this one |
|---|---|---|
| Language | TypeScript (Node 22 LTS) | The load-bearing artefact of this system is a set of input/output schemas. One zod definition gives the runtime validator, the static type shared between gate and Control Plane, and the JSON Schema handed to the model as a tool definition. A second language would mean maintaining that contract twice. |
| Datastore | PostgreSQL 16 + pgvector |
The mirror, the queue, the activity log and the guardrails are in one transactional boundary, so "commit once under a lease" is a real transaction rather than a distributed handshake. Semantic search over the proof and objection libraries sits in the same database instead of a second one. |
| Queue | Postgres SELECT … FOR UPDATE SKIP LOCKED |
Peak volume is thousands of intents a day, not millions. A dedicated broker would add a second durability story and a second failure mode to buy throughput nobody needs. Revisit if sustained queue depth crosses five figures. |
| Model access | Provider-agnostic router; Claude as the default binding | Agents declare a task class; a router resolves it to a provider and model that DMA chooses and can change without a deploy. Hosted, cloud-tenancy and self-hosted models are all bindable. See below — the vendor is configuration, not architecture. |
| Agent runtime | Anthropic SDK direct, no agent framework | The guardrails have to be inspectable and unit-testable. A framework that owns the loop also owns the place where a rule would have to live, and puts an abstraction between a reviewer and the thing being reviewed. The loop here is roughly forty lines. |
| Control Plane UI | Next.js (App Router), server components | Eight internal screens, single-digit concurrent users, no public surface. The least load-bearing choice in the table — swap it freely. |
| ESP with dedicated sending domain + inbound parse webhook | Deliverability, warm-up, bounce and complaint feedback loops are the product being bought. Inbound replies arrive as a webhook, not an IMAP poll. | |
| Hosting | Containers on Fly.io, managed Postgres | Matches how RSM already ships. Nothing in the design depends on it; the only real requirements are a scheduler, a private network and a daily backup with a tested restore. |
| Auth | Google Workspace SSO, roles in Postgres | DMA already has the identity provider. Approver identity has to be a real named person for the four-eyes rule in section 05 to mean anything. |
The model router
Agents do not name a model. They declare a task class, and a router resolves that class to whatever provider and model DMA has bound to it. The split into classes is not a cost trick — it is that reading 30,000 records and writing an email a prospect will read are different jobs, and the second is where quality converts. Which engine does each job is a separate question, and it is DMA's to answer.
The three classes below are a property of the work. The binding underneath them is configuration: a hosted frontier model, a different vendor, a model on DMA's own cloud tenancy, or something self-hosted on DMA's hardware. Changing a binding is a policy change in the Control Plane, not a deploy.
| Task class | The work | What the model must do | Agents |
|---|---|---|---|
bulk |
High-volume reading and classification over records, transcripts and replies. | Cheap per token and adequate at long context. Quality bar is modest; volume is not. A reduced-cost asynchronous lane is worth real money here and is the main reason the one-time 30K pass is inexpensive. | A04, A08, A11 |
judgement |
Structured extraction and bounded selection where the answer has to be defensible afterwards. | Reliable adherence to a supplied schema and close instruction-following. Prose quality matters less than never inventing a field or drifting off the enum. | A02, A03, A05, A06, A10, A14, A15 |
authored |
Anything a prospect will read, and anything commercial. | The best available quality on customer-facing writing and negotiation, plus reasoning depth. This is the class where paying more is justified, because the output is the product. | A01, A07, A09, A12, A13 |
An example binding
What follows is the binding I would start with and the one the cost estimates in Part Two, section 05 were computed against. It is an illustration of how a binding is expressed, not a recommendation you are locked into — every row is a value in a config table.
| Class | Example binding | List price / MTok | Equally valid alternatives |
|---|---|---|---|
bulk | claude-haiku-4-5 | $1 in · $5 out | Any small hosted model, or a self-hosted open-weights model on DMA hardware — this class is the best candidate for local inference, since the work is high-volume, low-stakes and never customer-facing. |
judgement | claude-sonnet-5 | $2 in · $10 out | Any mid-tier hosted model with dependable schema adherence. Worth benchmarking two candidates against the evaluation set in section 14 before fixing one. |
authored | claude-opus-5 | $5 in · $25 out | A frontier model from any vendor. The class where I would be slowest to economise, and the one where a measured comparison on DMA's own copy is worth running. |
export type TaskClass = "bulk" | "judgement" | "authored";
// A binding is a row in setting_version (scope = 'model_binding'), so changing
// one is effective-dated, needs a second approver, and is replayable — a run
// from March can be re-read against the model that actually produced it.
export interface ModelBinding {
provider: string; // 'anthropic' | 'openai' | 'google' | 'bedrock' | 'vertex'
// | 'azure' | 'ollama' | 'vllm' | any adapter we write
model: string;
endpoint?: string; // self-hosted or gateway; omitted for first-party APIs
inputPerMTok: number; // feeds the cost ledger; 0 for owned hardware
outputPerMTok: number;
capabilities: {
nativeSchema: boolean; // constrained decoding, or we shim it (below)
asyncBatchLane: boolean; // reduced-cost bulk submission
systemChannel: boolean; // operator instructions separable from user content
promptCache: boolean;
};
}
// Every provider implements one interface. Nothing above this line knows the vendor.
export interface Provider {
run<T>(req: {
binding: ModelBinding;
schema: ZodType<T>;
system: Message[]; // policy. never contains untrusted text.
input: Message[]; // content. always treated as data.
tools?: ToolDef[];
}): Promise<Run<T>>;
}
// Agents ask for a class, never a model.
const run = await router.run("authored", { schema: ProposalShape, system, input });
Four capabilities matter. Two are required, and two degrade gracefully — which is what makes swapping a binding a real option rather than a claim.
- Schema-constrained output — required, shimmed where absent. Every agent returns a validated object; a response that fails validation is a failed run, never a partial write. Where a provider constrains decoding natively we use it; where it does not, the adapter validates and retries, and agents above the boundary see no difference beyond a slightly higher failure rate on that binding.
- A separate operator channel — required, and the one to check first. Section 13 rests on operating instructions arriving on a different channel from visitor text. Where a provider offers no such separation mid-conversation, containment leans harder on the tool grant and on output validation, and I would not put that binding on
authoredfor the public agent. This is the capability worth testing before choosing, not after. - A reduced-cost asynchronous lane — optional. Without it the one-time 30K pass costs roughly twice the batched figure in Part Two, section 05. Nothing breaks; a line item moves.
- Prompt caching — optional. Without it the recurring bill rises and
run.cache_read_tokensstays at zero. Section 14 monitors it either way.
Three reasons, in the order they are likely to matter to DMA. The market moves faster than this project will: a binding that is right at kickoff may not be right at Phase 2, and the cost of being wrong should be a config change. Some classes have a genuine case for local inference — bulk reads 30,000 of DMA's own records and never speaks to a customer, so running it on DMA hardware is a defensible answer to both cost and data-residency questions. And a compliance position can force the issue: if the data processing agreement in section 15 lands somewhere restrictive, the binding is the thing that changes, not the architecture.
Six processes, one database
Every process below is stateless and horizontally scalable except the gate worker, which is deliberately single-flight per record. They share one Postgres and speak to each other only through it.
custom_id because batch results return in arbitrary order.Each service connects as its own Postgres role. svc_web cannot read setting_version, activity_log or deal_transaction; svc_agents has no write grant on ns_object. This is what makes the containment claims in section 13 structural rather than aspirational — a flaw in the public chat surface reaches exactly the rows its role was granted.
A local copy of a system we do not control
The mirror exists so that scoring, hygiene and planning can read 30,000 records repeatedly without hammering an API we have no rate-limit budget for. It is a cache, never an authority: Nutshell is the system of record and the mirror is always the thing that is wrong.
One table with a discriminator rather than seven typed ones, because the shape of a Nutshell object is not ours to control and DMA will add custom fields. Typed access happens through generated columns and views, so adding a field is a migration on our projection rather than on the store.
-- Nutshell objects, stored as fetched. Never edited locally.
create table ns_object (
kind text not null check (kind in ('account','contact','lead','activity','note','task','user')),
nutshell_id text not null,
payload jsonb not null,
content_hash text not null, -- sha256 of canonicalised payload
fetched_at timestamptz not null,
first_seen_at timestamptz not null default now(),
deleted_at timestamptz,
primary key (kind, nutshell_id)
);
create index on ns_object using gin (payload jsonb_path_ops);
create index on ns_object (kind, fetched_at); -- drives the staleness sweep
-- Change signals. Deliberately carries no payload: a notification says a record
-- is worth looking at, never what it now says. See section 12.
create table ns_change (
id bigserial primary key,
kind text not null,
nutshell_id text not null,
source text not null check (source in ('webhook','poll','reconcile')),
event_id text unique, -- replayed deliveries collapse here
seen_at timestamptz not null default now(),
handled_at timestamptz
);
-- Reconciliation findings, by field class. Section 12 explains the two kinds.
create table drift (
id bigserial primary key,
kind text not null,
nutshell_id text not null,
field text not null,
field_class text not null, -- 'contact_detail' | 'stage' | 'value' | 'ownership' | 'ai'
mirror_value jsonb,
live_value jsonb,
verdict text not null check (verdict in ('ours_superseded','unaccounted')),
found_at timestamptz not null default now(),
resolved_at timestamptz
);
content_hash earns its column
It does three jobs. It tells the bulk scheduler whether a record needs re-reading, which is the whole reason the recurring model spend is a fraction of the one-time spend. It gives the reconciliation pass a cheap equality test. And it is the fingerprint the adapter writes into created objects so a blind retry can be recognised as a duplicate — the exactly-once mechanism in section 12.
Facts with provenance, not a summary blob
The Canonical asks for an Account Digital Twin: one living picture of an account that every part of the system reads from. The tempting implementation is a generated paragraph per account. That fails the Canonical's own explainability rule the first time somebody asks how do you know that, and it decays silently because nothing can be individually corrected or expired.
So the twin is not a document. It is a set of atomic facts, each carrying the sentence it came from, the object that sentence lives in, and the run that extracted it. Superseding a fact is a row, not a rewrite.
create table twin_fact (
id bigserial primary key,
account_id text not null,
kind text not null, -- 'incumbent' | 'budget_signal' | 'service_fit' | 'objection'
-- | 'commitment' | 'contact_role' | 'timing' | 'lost_reason'
value jsonb not null,
confidence numeric(3,2) not null check (confidence between 0 and 1),
source_kind text not null, -- 'note' | 'activity' | 'transcript' | 'chat' | 'enrichment'
source_ref text not null, -- ns_object key, or transcript id
source_quote text not null, -- the sentence. no quote, no fact.
extracted_by uuid not null references run(id),
valid_from timestamptz not null default now(),
superseded_by bigint references twin_fact(id),
unique (account_id, kind, source_ref, md5(value::text))
);
create index on twin_fact (account_id, kind) where superseded_by is null;
-- Derived, recomputed by code from the facts above. Never written by a model.
create table account_score (
account_id text primary key,
score smallint not null check (score between 0 and 100),
tier char(1) not null check (tier in ('A','B','C','S')), -- S = suppress
features jsonb not null, -- the extracted inputs, for replay
weights_version text not null, -- which approved weight set produced this
computed_at timestamptz not null default now()
);
A model reads the prose and emits features; code turns features into a number using an approved, versioned weight set. Re-running the arithmetic on stored features reproduces any score exactly, and changing the weights re-scores 30,000 records without spending a token. The part that cannot be reproduced — the reading of a nine-year-old call note — is isolated in one column you can inspect.
Queue, log, policy, state
Four groups of tables. The intent queue is what agents produce. The activity log is what the gate leaves behind. The policy tables are what DMA edits. The transaction table joins three external systems that each know a third of the truth.
Intents — the only thing an agent can produce
create type intent_status as enum ('queued','checking','held','committed','discarded','failed');
create table intent (
id uuid primary key default gen_random_uuid(),
type text not null, -- 'nutshell.write' | 'nutshell.task' | 'email.send'
-- | 'proposal.send' | 'payment.request' | 'field.restore'
account_id text,
lead_id text,
payload jsonb not null, -- validated against the schema for `type`
snapshot jsonb not null, -- state the agent computed from
material_fields text[] not null, -- what invalidates this intent (gate step 2)
proposed_by uuid not null references run(id),
idempotency_key text not null unique,
batch_id uuid references send_batch(id),
status intent_status not null default 'queued',
not_before timestamptz not null default now(), -- the hold window lives here
attempts smallint not null default 0,
last_error text,
created_at timestamptz not null default now()
);
create index on intent (not_before) where status = 'queued';
-- One machine action in flight per opportunity. The gate takes this lock;
-- nothing else in the system may write to a lead without holding it.
-- pg_advisory_xact_lock(hashtextextended(lead_id, 0))
The activity log — audit and undo are one table
Nutshell does not retain prior field values, so they have to be retained here. The gate re-reads the live record immediately before it writes, so the outgoing value is already in hand; it travels with the entry. Restoring a field means re-applying the prior value from its most recent entry — and the restore goes back out as an ordinary intent, so it re-checks live state and is itself logged.
create table activity_log (
id bigserial primary key,
at timestamptz not null default now(),
run_id uuid references run(id),
intent_id uuid references intent(id),
action text not null,
target_kind text, target_id text, field text,
prior_value jsonb, -- read live at gate step 1; null when the action created something
new_value jsonb,
rule_id text, -- the setting_version that permitted it
reversible boolean not null,
reversed_by bigint references activity_log(id),
cost_usd numeric(10,6)
);
create index on activity_log (target_kind, target_id, field, at desc);
-- Restore = the most recent prior value for this field, re-proposed through the gate.
create view restorable as
select distinct on (target_kind, target_id, field)
target_kind, target_id, field, prior_value, at, run_id
from activity_log
where reversible and reversed_by is null and field is not null
order by target_kind, target_id, field, at desc;
Policy — the tables DMA edits
-- Every load-bearing setting, effective-dated and signed off by a second person.
create table setting_version (
id bigserial primary key,
scope text not null, -- 'guardrail' | 'weights' | 'routing' | 'field_permission'
-- | 'prompt' | 'playbook' | 'rate_card' | 'model_binding'
key text not null,
value jsonb not null,
proposed_by text not null,
approved_by text,
approved_at timestamptz,
effective_from timestamptz not null,
effective_to timestamptz,
constraint four_eyes check (approved_by is null or approved_by <> proposed_by),
constraint inert_until_approved check (approved_at is not null or effective_from > 'infinity'::timestamptz)
);
create index on setting_version (scope, key, effective_from desc);
-- Approved precedent. The key is a hash of a fixed attribute tuple — a lookup,
-- never a judgement. Seeded in Discovery from closed-won history.
create table precedent_key (
key_hash text primary key,
attributes jsonb not null, -- service_mix, size_band, term_months, payment_terms,
-- template_version, discount_rung, scope_flags
source text not null check (source in ('historical','escalation')),
approved_by text not null,
approved_at timestamptz not null,
retired_at timestamptz
);
-- Where each of the seven triggers is sent. A mapping, not a process:
-- no timers, no backups, no coverage rules. Part Two, section 09.
create table escalation_route (
trigger smallint primary key check (trigger between 1 and 7),
destinations jsonb not null, -- [{kind:'slack_channel'|'nutshell_user'|'email', id:'…'}]
updated_by text not null,
updated_at timestamptz not null default now(),
constraint never_empty check (jsonb_array_length(destinations) > 0)
);
-- The approved list of permissions by field. DMA supplies this; the gate enforces it.
create table field_permission (
object_kind text not null,
field text not null,
may_write boolean not null,
autonomy text not null check (autonomy in ('write','approve','propose')),
on_collision text not null check (on_collision in ('human_wins','system_wins','show_both')),
lock_days smallint, -- null = until the next substantive human touch
primary key (object_kind, field)
);
Contact state, runs, and the transaction join
-- Suppression is computed, not remembered. Nothing has to set a flag correctly.
create table contact_state (
contact_id text primary key,
consent_basis text,
verify_result text check (verify_result in ('valid','invalid','risky','unknown')),
verified_at timestamptz,
unsubscribed_at timestamptz,
bounced_at timestamptz,
complained_at timestamptz,
last_sent_at timestamptz,
sends_30d smallint not null default 0,
suppressed text generated always as (
case when unsubscribed_at is not null then 'unsubscribed'
when complained_at is not null then 'complained'
when bounced_at is not null then 'bounced'
when consent_basis is null then 'no_consent_basis'
when verify_result in ('invalid','risky') then 'unverified'
else null end) stored
);
-- One row per agent invocation. Traceability and cost in the same place.
create table run (
id uuid primary key default gen_random_uuid(),
agent_id text not null, -- 'A09'
agent_version text not null, -- setting_version id of its prompt
task_class text not null, -- 'bulk' | 'judgement' | 'authored'
provider text not null, -- resolved binding, so cost can be split by vendor
model text not null,
input_digest text not null,
outcome text check (outcome in ('ok','schema_fail','refusal','timeout','error')),
output jsonb,
input_tokens int, output_tokens int, cache_read_tokens int,
cost_usd numeric(10,6),
started_at timestamptz not null default now(),
ended_at timestamptz
);
-- Three providers each know a third of the truth. This is the join.
create table deal_transaction (
lead_id text primary key,
state text not null check (state in (
'assembling','held','sent','negotiating',
'signed','signed_unpaid','paid','lost','refunded')),
artefact_hash text, -- exactly what was approved; re-checked at send
requested_by text, approved_by text,
esign_id text, esign_state text,
payment_id text, payment_state text,
expected_amount numeric(12,2),
collected_amount numeric(12,2),
updated_at timestamptz not null default now(),
constraint separation check (approved_by is null or approved_by <> requested_by),
constraint paid_means_paid check (state <> 'paid' or collected_amount = expected_amount)
);
four_eyes means nobody can approve their own policy change, including me. separation means the person who requested a binding document is not the person who released it. paid_means_paid means a signature alone cannot close a deal and neither can a payment for the wrong figure — a bug in the payment handler produces a failed transaction, not a false Closed Won. Rules enforced here cannot be forgotten by a future code path.
The Nutshell-side fields
The custom fields written back to Nutshell are set out in full in Part Two, section 03 and are not repeated here. Two implementation notes: every one is governed by a field_permission row before the gate will write it, and ai_run_id carries the run.id that produced the values, so any field a rep is looking at opens the activity log entry behind it.
Fifteen agents, each with one job
Each agent is a prompt version, an input schema, an output schema, a tool grant and a task class — never a named model. None of them writes anywhere; all of them emit intents or facts. L0–L4 are the Canonical's own autonomy levels. The one agent that talks to strangers is marked, because it is the one to review hardest.
content_hash moves.
bulk·L0·batch
Pricing. Tier boundaries. Consent and suppression. Payment retry and dunning. Whether there is approved precedent. Whether to send. Each is a deterministic rule in the gate, and each one exists as a rule because the failure mode of getting it wrong is a bad price, an unlawful send, a double charge or an unapproved contract. A model proposing is fine; a model deciding is not.
A schema, a grant, and a run record
Every agent is the same three things. The example below is A12, because it carries the clearest version of the central idea: the model cannot emit a price, because the schema has nowhere to put one.
import { z } from "zod";
import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod";
// The commercial shape of a deal — services, scope, term. Not money.
export const ProposalShape = z.object({
services: z.array(z.object({
service_code: z.enum(RATE_CARD_CODES), // closed set, from the approved rate card
quantity: z.number().int().positive(),
scope_notes: z.string().max(400),
})).min(1),
term_months: z.union([z.literal(6), z.literal(12), z.literal(24)]),
payment_terms: z.enum(APPROVED_PAYMENT_TERMS),
discount_rung: z.number().int().min(0).max(4), // a rung, never a percentage
justification: z.string().max(600),
evidence: z.array(z.object({
fact_id: z.number(), quote: z.string(),
})).min(1), // no evidence, no proposal
});
// There is no `price`, `total`, `discount_pct` or `monthly` field, and adding one
// is a reviewed change to this file — not something a prompt can talk its way into.
// A12 asks for a class. Which provider and model answers is DMA's binding,
// resolved at call time and recorded on the run for replay.
const res = await router.run("authored", {
schema: ProposalShape,
system: [A12_SYSTEM], // policy channel — no untrusted text, ever
input: [renderTwin(twin)], // content channel — always data
});
if (!res.ok) throw new SchemaFailure(run.id); // a failed run, never a partial write
const priced = priceFromRateCard(res.value, rateCardAt(now)); // code owns the money
await proposeIntent({ type: "proposal.send", payload: priced, proposedBy: run.id });
Tool grants are configuration, not conversation
// The public agent's entire reach. Attached to this agent version at deploy time. export const A01_TOOLS = [ searchApprovedAnswers, // published content only fetchApprovedCaseStudy, requestSiteReview, // enqueues A02; returns an acknowledgement, not a result checkAvailability, bookMeeting, captureLead, // creates an intent. does not write to Nutshell. askForHuman, // trigger 4. no threshold, no deferral. ] as const; // Absent from that array, and therefore unreachable no matter what a visitor writes: // any pricing tool · any write to an existing Nutshell record · any read of // another account · the guardrail registry · the activity log · e-sign · payments. // There is no sequence of messages that adds an entry to a frozen array.
An agent's system prompt is a setting_version with scope = 'prompt', so it inherits the four-eyes constraint and the effective-dating for free: a prompt change needs a second approver, and a run from March replays against March's prompt. The prompt text lives in the repository and is loaded by version — the row is the authorisation record, not the storage.
The nine checks, in order, every time
The mechanism is set out in full in Part Two, section 07. What follows is its implementation: the loop, the lock, and where each check reads from.
async function claimAndCommit() {
await db.tx(async (tx) => {
const intent = await tx.one(`
select * from intent
where status = 'queued' and not_before <= now()
order by not_before
for update skip locked limit 1`);
if (!intent) return;
// Single-flight per opportunity. Held for the whole transaction, so two
// agents can never act on the same lead concurrently.
await tx.query(`select pg_advisory_xact_lock(hashtextextended($1, 0))`, [intent.lead_id]);
const live = await nutshell.read(intent.lead_id); // 1 live state, not the snapshot
if (materiallyChanged(intent, live)) return discard(tx, intent, "stale"); // 2
if (await humanTouched(tx, live, intent)) return replan(tx, intent, "human"); // 3
const block = await autonomyAndSuppression(tx, intent, live); // 4
if (block) return hold(tx, intent, block);
const bounds = await guardrails(tx, intent); // 5
if (!bounds.ok) return escalate(tx, intent, bounds.trigger);
const trig = await escalationTriggers(tx, intent, live); // 6
if (trig) return escalate(tx, intent, trig);
if (isOutbound(intent) && !(await claimsResolve(tx, intent))) // 7
return hold(tx, intent, "unresolved_claim");
const prior = extractPriorValues(live, intent); // captured before we overwrite
const result = await adapter.commitOnce(intent); // 8 idempotent, rate-limited
await writeActivityLog(tx, intent, prior, result); // 9 audit and undo, one row
});
}
| Check | Reads from | Exit |
|---|---|---|
| 1 · Live re-read | Nutshell API, not the mirror | Retry |
| 2 · Materiality | intent.material_fields vs. live | Discard |
| 3 · Human priority | ns_object activity + field_permission.on_collision | Discard |
| 4 · Autonomy & suppression | contact_state.suppressed, ai_autonomy, kill switches, circuit breakers | Hold |
| 5 · Guardrails | setting_version effective at now() | Escalate |
| 6 · Triggers 1–7 | precedent_key, deal size, terms allow-list | Escalate |
| 7 · Claims | A15 resolutions vs. twin_fact + proof library | Hold |
| 8 · Commit once | Adapter, under the advisory lock | Retry or no-op |
| 9 · Log, then hold | — | — |
There is no separate scheduler for it. A reversible action is committed to the log and its follow-on effect is queued with not_before = now() + window, so cancelling is a row update inside a transaction rather than a race against a timer. Outbound gets a longer window than a CRM field update, because a CRM write can genuinely be put back and an email cannot.
Visitor to qualified opportunity
Two things are worth reading off the diagram. Untrusted input enters at two points, not one — the visitor's messages and the pages A02 fetches — and neither reaches a tool. And every arrow that leaves the agent row passes through the gate; there is no path from a conversation to Nutshell that skips it.
Qualified to collected, including when it breaks
The happy path is five states. The value is in the three branches below it, and in the rule that only a payment matching the expected amount moves a deal to Closed Won.
paid_means_paid enforces that in the schema. The three red states are the ones that exist so a deal can never quietly disappear between two providers.| Rule | Implementation |
|---|---|
| What was approved is what goes out | deal_transaction.artefact_hash is set at approval and re-checked immediately before send. A regenerated draft or an edited rate card underneath it fails the comparison and returns the item for approval. |
| Requester ≠ releaser | The separation constraint. Not a code path that can be missed. |
| Verify provider callbacks | Signature verification on every e-sign and payment webhook, plus a sweep that asks each provider for the current state of anything open longer than its expected window. A callback that never arrives is normal, not exceptional. |
| Card data never lands here | Processor-hosted checkout only. No card field exists in any schema in this document. |
Thirty thousand records to a booked meeting
This is the engine that reaches revenue first, and it is mostly batch work. The only human step in the steady state is releasing a send batch — and that is an operator action taken once per batch, not a per-record approval.
Holdout records are scored, tiered, given a why-now line and drafted exactly like the rest — they are only withheld at the point of sending. That costs a little model spend and buys the only honest comparison available: the control group differs from the treated group in one variable, whether a message went out, rather than in how much attention the system paid it.
Reading, writing, and disagreeing nightly
The adapter is the only module that knows Nutshell exists. Everything above it works in our own types, which is what makes the largest unknown in the project survivable: if the API turns out to be polling-only or field-restricted, the blast radius is this module plus the sync service.
What I have been able to confirm without the account
| Question | What the public documentation says | Consequence here |
|---|---|---|
| Can Nutshell tell us what a field held before it changed? | No. Webhook payloads carry the entity's current state with an empty changes array, and Nutshell's own guidance is to treat a webhook as a prompt that something changed, then re-read. |
Prior values must be retained on our side — activity_log.prior_value. It also independently confirms the "signals, never data" rule. |
| Is there a usable audit trail we could lean on? | The Audit Log is an Enterprise-plan admin screen covering logins, bulk edits and exports. It cannot be exported and has no API. | Not available as a mechanism. Our own log is not duplication. |
| Is there a change feed at all? | Yes — an events feed and webhooks, plus a deletion-events endpoint. | Webhook-driven sync is the primary path; polling is the fallback, and the design works either way because neither is trusted for content. |
| Per-write idempotency support? | Unverified Not documented publicly. | The adapter carries its own fingerprint scheme below. If native support exists, we use it and delete that code. |
| Rate limits under a 30K backfill? | Unverified | Sizes the backfill window. Section 15. |
Writing exactly once
The failure that matters is narrow: we send a write, the connection dies before the response arrives, and we do not know whether it landed. Blind retry gives a prospect two identical notes.
async function commitOnce(intent: Intent): Promise<CommitResult> {
const fp = fingerprint(intent); // deterministic: content + run id
if (intent.creates) {
// Created objects carry their fingerprint in a field we own.
const existing = await findByFingerprint(intent.creates.kind, fp);
if (existing) return { status: "noop", id: existing.id }; // the first attempt did land
return await create({ ...intent.payload, ai_fingerprint: fp });
}
// Field updates are easier: if it already holds what we meant to write, we are done.
const live = await read(intent.target);
if (equal(live[intent.field], intent.payload.value)) return { status: "noop" };
return await update(intent.target, { [intent.field]: intent.payload.value });
}
The nightly disagreement
A full pass compares the mirror against Nutshell nightly; a rolling sample checks a slice throughout the day, because a bad deploy at 9am should not have until midnight to do damage. Findings sort into two kinds, and they matter differently.
| Verdict | Meaning | Action |
|---|---|---|
ours_superseded | A value we wrote no longer matches — a person changed it and we missed the notification. | Correct the mirror; the field becomes theirs under field_permission.on_collision, exactly as if we had seen the edit. |
unaccounted | Neither side can explain the value. | An integrity problem. This is the one that should make somebody nervous. |
Thresholds are per class, not per record, because a percent of contact-detail churn overnight is ordinary and the same drift in deal stages or values is not. When a class crosses its threshold, writes for that class pause automatically and the queue holds rather than drains — nothing is dropped, and somebody has to look before it resumes.
Four mechanisms, none of them a prompt
Part Two, section 08 makes the argument: the question is not whether the agent refuses, but what follows if it does not. These are the four places that answer is enforced.
A01_TOOLS, and there is no message sequence that appends an eighth.
svc_web has no grant on the guardrail, activity-log or transaction tables, and no write grant on the mirror. A compromise of the public surface reaches what the role was granted and nothing else.
→ A flaw in public chat does not become read access to 30,000 records, because that process was never able to read them.
Ceilings, caches, and what gets measured
Part Two, section 05 puts numbers on the spend. This is how those numbers are held in place, and what is instrumented so the first month's real figures can replace the estimates.
content_hash, so unchanged records and recently reviewed sites are not re-read. Cache effectiveness is monitored on run.cache_read_tokens; a sustained zero means something volatile crept into the prefix.run row carries its cost, and every activity_log row links to its run. That makes cost-per-booked-meeting and cost-per-closed-deal ordinary queries rather than a reporting project — and it is the only way to tell whether a cheaper model on a given agent actually saved anything.Testing, and the part that is hard to test
| Layer | How |
|---|---|
| Guardrails and triggers | Ordinary unit tests. They are code, so the assertion "will not discount past the floor" is a test that either passes or fails. This is the main argument for the whole architecture. |
| The gate's concurrency | Integration tests that fire an agent intent, a simulated rep edit and a webhook at the same record simultaneously, and assert on which survives. The hardest part of the system deserves the most adversarial test. |
| Adapter against Nutshell | Contract tests against a sandbox account, plus a recorded-fixture suite. A stand-in CRM proves the software works against our assumptions, which is a different claim from working against Nutshell — section 15. |
| Agent output quality | A fixed evaluation set of DMA examples per agent, scored on schema validity, claim resolution and — for customer-facing agents — human rating. Prompt changes run it before the second approver sees them. |
| Deliverability | Seed-list sends and inbox placement checks before the first real batch, then continuous monitoring. Not testable after the fact. |
Three: local against fixtures, staging against a Nutshell sandbox with an e-sign sandbox and payment test mode, and production. Staging never holds real contact records, and the ESP credential in staging points at a sink. The only way to send a real email from this system is to be in production and to pass every check in section 08.
What Discovery has to settle before this is buildable
These are the technical questions this document could not answer from outside. Each one changes code rather than adding to it, which is why they belong at the front of the project rather than in week three.
| Decision | What it changes | How it gets settled |
|---|---|---|
| Nutshell write surface and rate limits | Highest Whether the sync design is webhook-driven or polling-driven, how long the backfill takes, and whether a state-reconciliation layer no estimate currently includes is needed. | A read-only credential and two days. This is the one question no prototype can answer, including mine. |
| Custom field capabilities on the live account | Whether the field schema in Part Two, section 03 is writable as designed, or has to be folded into fewer fields. | Same credential, same two days. |
| Bulk-edit-from-list-view support | Decides the "Approve All" surface — email digest with signed links, Nutshell bulk edit, or one Control Plane screen. | Testable against the live account. Part Two, section 02 sets out the three answers. |
| Website platform and chat widget | Whether A01 ships as an embedded widget we host or inside an existing chat product, and how live human takeover works. | DMA decision plus a look at the site. |
| Triggers 1 and 7 | Whether a deal above every previously approved size proceeds or escalates. Currently unresolvable in code because the Canonical says both. | A commercial call from DMA. Needed before the transaction cluster ships. |
| Rate card structure | The shape of RATE_CARD_CODES and the discount ladder — the two closed sets that make A12's no-price-field design work. |
DMA's written pricing policy. The single most useful thing to start on today. |
| Consent basis per jurisdiction | How much of the 30,000 is contactable, which is the core value thesis; and whether contact_state.consent_basis needs a per-jurisdiction dimension. |
DMA's legal position. Not a resolution from me. |
| Model provider and data processing | Which binding sits behind each task class, and whether any of them may hold DMA’s records and transcripts. A restrictive answer moves bulk to DMA hardware or a cloud tenancy; it does not change anything above the router. |
DMA’s choice, with a data processing agreement per retained provider. Worth benchmarking two candidates per class on DMA’s own examples before fixing one. |
| Transcript retention scope | Whether A10 reads existing Zoom transcripts only, or new recordings — the Canonical's one High-risk row. | Legal review, per the Canonical's own gate. |
The adapter and the gate, in that order, before any agent. They are the two pieces every other piece depends on, they are the two most likely to be quietly wrong in a system that demos well, and they are the only parts whose design changes materially once the Nutshell questions above are answered. Everything in section 06 is comparatively easy — and comparatively easy to replace.
No live data, and no verified Nutshell account behind any of it — see section 15.
What would settle this
Everything above is a design I can defend, and none of it is verified. That distinction carries more weight here than it would in most specs, because the gap is narrow and closable: a read-only Nutshell credential and two days would turn the largest unknown in this document into a written answer. Whether the write surface, the custom field types and the rate limits support what is designed above is a measurement rather than a judgement, and it is the only open item whose answer can invalidate work already done instead of simply adding to it.
It is also the reason this arrived as three documents rather than as something you could click through. A working prototype would show that the software behaves as its author intended, against a CRM its author chose. It would answer none of what those two days would answer, and it would look considerably more finished while telling you less. Part Two, section 00 makes that case properly.