Robot Sea Monster Games Implementation Approach Pre-Discovery Draft

DMA AI Revenue System

An approach to the v3 Canonical: how DMA sells today as I currently read it, where a person stays in control once a machine is running it, and what has to be built for that to hold.

Prepared for Jeev Trika & Josh Moody Against DMA AI Revenue System v3 Canonical Author Ryan Abrams
Part One

The workflow, as I understand it today

One inbound message, followed all the way through — and then the same mechanism laid out across the whole journey. This is DMA's motion as I currently read it, not a process I am proposing you adopt.

00 / The workflow

What happens when somebody gets in touch

Start with one message from one visitor and follow it. Almost everything in this document is visible in that single turn: how the message is handled, what the agent can and cannot reach, when qualification happens, and what has to be true before anything is written down or sent.

What this is, and what it is not

The top row is DMA's process as I currently understand it — assembled from the v3 Canonical and our two calls, and reported rather than proposed. I have no reason to think I understand how DMA sells better than DMA does, and the parts I am least sure of are the ones nobody has written down. Where this drawing and your actual practice disagree, your practice is right and the drawing is the thing that is wrong. Correcting it is the first hour of Discovery, and it is cheap now and expensive later.

Everything below the top row is mine — how I would build the machinery so that those stages happen reliably, and the same way every time. So there are two different kinds of disagreement available in these pictures, and both are useful: argue with the top row on the facts, and with everything under it on the engineering.

The point of automating a sales motion is to make what already works happen more often. A system built on a process the team does not recognise is worse than no system, which is why this page comes first.

THE VISITOR’S TURN — synchronous, start to finish in seconds → Visitor message arrives on the content channel, as quoted data Context assembled policy channel, history, approved answers A01 runs tools fixed before the turn begins Reply drafted not shown to anyone yet Outbound check every figure resolves, or the reply is held Visitor sees it one turn, 2–4 seconds WHAT A01 CAN REACH DURING THE TURN — its entire surface · search approved answers · fetch an approved case study · check availability · book a meeting · capture the lead → intent · request a site review → A02 · ask for a human → trigger 4 not in the list, so unreachable: pricing · Nutshell writes · other accounts · payments A02 Site Reviewer fetches the prospect’s public pages returns labelled data, never instructions WHAT THE SAME MESSAGE SETS OFF — asynchronous, after every visitor turn → Transcript so far re-read every turn A03 Qualification services, scale, timing, authority, budget twin_fact rows each with the sentence it came from Scorer code, versioned weights score + tier 0–100 · A / B / C / S COMMIT GATE nine checks the same message, read again for facts Nutshell lead + ai_* fields written Slack hot-lead alert · SLA clock starts Escalation route trigger named · AI holds the thread
Fig. 0 — One message, end to end, on the flow as currently understood. The top lane is what the visitor experiences and it finishes in seconds. The bottom lane is what the same message sets off for DMA, and it finishes when the gate says so. Nothing crosses from the bottom lane into Nutshell, Slack or a person without passing the nine checks.

Three things worth reading off it

The visitor's words and the operating instructions never travel together. The message arrives on the content channel as quoted data; the policy — who the agent is, what it may do — arrives separately and cannot be amended by anything a visitor types. That separation, plus a tool list fixed before the turn starts, is the whole of the containment story: the agent's reach is the seven tools in that box and nothing else, no matter how persuasive the message.

Qualification is not a stage somebody declares. It runs after every single turn. The transcript is re-read, A03 pulls out services, scale, timing, authority and budget signal — each one stored with the sentence it came from — and a scorer written in code, not a model, turns those facts into a number and a tier. Nobody has to notice that a visitor became interesting; the threshold does, on the turn it becomes true.

The fast path and the committing path are deliberately different. A reply is checked for unresolved claims and shown in seconds, because a prospect will not wait. Anything that changes state — a Nutshell write, a hot-lead alert, a booking, a hand-off — becomes an intent and queues for the gate, where it is re-validated against live state before it commits. The visitor is never waiting on the gate, and the gate is never rushed by the visitor.

The same shape, across the whole journey

That loop is not special to first contact. It repeats at every stage: something arrives, an agent proposes, the gate decides, a person sees the result. What changes from stage to stage is which agents are involved and how much the gate is willing to let through on its own. The stages themselves are the part to check against how DMA actually sells — if a step is missing, out of order, or never happens the way it is drawn here, that is worth saying before anything is built on it.

WHAT THE PROSPECT EXPERIENCES → ENGINE 1 OR 2 First contact a website visitor, or a dormant record we reached CHAT OR E-MAIL Conversation questions answered, objections handled SCORED Qualified their site reviewed, their situation understood WITH A PERSON Meeting booked into a real calendar, then held WITHIN GUARDRAILS Proposal priced from the rate card, sent to sign CLOSED WON Signed & paid deposit captured, they become a client THE COMMIT GATE nothing above this line happens until it has passed through here · every Nutshell write, every e-mail, every proposal, every charge THE AGENTS THAT DO THE WORK — each under the stage it serves A01 Site Conversation — runs the whole conversation A07 Reactivation Drafter A08 Reply Classifier A02 Site Reviewer A14 Methodology Selector A12 Proposal Assembler A09 Reply Responder A03 Qualification Extractor A10 Transcript Extractor A13 Negotiation Agent A15 Claims Resolver no agent here — this is bought, and gate-driven RUNNING CONTINUOUSLY UNDERNEATH — deciding who enters the journey at all over 30,000 Nutshell records A04 Record Reader A05 Why-Now Writer A06 Next Best Action A11 Duplicate Candidate
Fig. 1 — The journey as a map rather than a flow: six stages as I currently understand DMA's process, and and which agents serve each one. The arrows run upward because that is the direction of proposal — agents suggest, the gate decides, the prospect sees the result. The one stage with no agent under it is the one where DMA buys the software rather than building it.

Three things to take from the map

The top row is the only part anyone outside DMA sees. Everything else — fifteen agents, a scoring pipeline, a mirrored database, a nightly reconciliation — exists to make those six boxes happen reliably. That is also why almost none of this system has a user interface, and why this arrived as a set of documents rather than as a demo: the parts that decide whether it works are not the parts a demo can show. Part Two, section 00 makes that case.

The gate is a single line every action crosses, not a step in a sequence. Drafting an e-mail, writing a CRM field, sending a proposal and charging a deposit are wildly different actions, and all four are re-validated, checked against DMA's rules and committed exactly once by the same piece of code. It is the first thing I would build and the thing most likely to be quietly missing from a system that demos well.

The bottom band is where the money is. Before a dormant record ever becomes a first contact, it has been read, scored, tiered and given a reason to be worth contacting today. Thirty thousand records are the fastest path to real revenue, which is why that work comes first in the build sequence rather than the website chat.

Part Two

Surface & control map

How the Canonical's flow maps onto Nutshell, onto services you already pay for, and onto the one piece of software that has to be purpose-built.

00 / Read this first

What this is, and what it deliberately is not

This is a surface and control map, not a proof of concept. It answers one question in detail: when the Canonical says the AI qualifies, scores, reaches out, negotiates, proposes, signs and collects — where does a human actually look, and what is the software that keeps all of it honest?

A generated proof of concept would have been the easier thing to send, and I would rather not lead with one. It shows a UI that has never touched your Nutshell account, priced against no pricing policy, with no idea whether your API supports the writes it assumes. It looks finished and it is unfalsifiable. It also fails the Canonical, which says plainly: no custom approval UI, Nutshell remains the single system of record.

So this document does the opposite. It takes your requirements as written, follows them to their consequences, and shows you the consequences — including the four or five places where the Canonical asks for two things that cannot both be true. Every screen illustrated here is a sketch of a real surface I would expect to build or configure, drawn to show placement and content, not visual design.

The one-sentence version

Almost none of this system has a user interface. The reactivation engine, the scoring pass, the reply handling, the write-back — all of it is background work that surfaces inside Nutshell as fields, tasks and lists your reps already know how to read. The custom software you actually need is small, and it is not a sales tool: it is the control plane that holds the guardrails, the action queue, the audit log and the transaction state, and reps will rarely open it.

On price and dates

I have not put a fixed price in this document, and the milestone dates below are shapes rather than commitments. That is a deliberate position rather than an evasion: the two largest cost variables in this project — whether the Nutshell API supports the writes the design assumes, and how much of the 92 Phase 1 rows constitutes the first release — are both unmeasured. Anyone quoting a firm number today is quoting across a gap they have not looked into. The sequence in section 11 is real; the calendar it lands on gets fixed once those two things are known. Section 05 does put numbers on the recurring model and vendor spend, which is a different question and an answerable one.

01 / The shape of it

Two engines, one commit gate, one system of record

The Canonical's architecture page describes "one AI brain." On a technical level that is a useful way to talk about it but not a thing you build. What you build is a large number of small, single-purpose agents — each one a prompt, an input schema and an output schema — plus a shared account memory they all read from, plus one piece of deterministic plumbing that every single one of them has to pass through before it is allowed to change anything in the real world.

That last piece is the whole design. It is what stops fifteen agents from stepping on each other, on your reps, and on Nutshell.

Engine 1 — Inbound Website AI salesperson Chat widget on the DMA site. Conversation, mini-consultant, objection handling, live re-scoring, booking.
Engine 2 — Reactivation 30,000 dormant records Audit, score, tier, next best action, staged personalised outreach, AI reply handling.
↓    ↓
Shared layer Account Digital Twin & agent library One living account memory — people, history, prior proposals, objections, commitments, current signals — written and read by every agent. Around it: scoring, extraction, drafting, classification and orchestration agents, each with a versioned prompt, a structured output contract, and the narrowest set of capabilities its job requires. Section 08.
↓
Everything passes through here Action Commit Gate & Control Plane No agent writes to Nutshell, sends an email, prices a deal or charges a card directly. Every action is proposed as an intent, re-validated against live state, checked against deterministic guardrails and the seven escalation triggers, then committed once — and logged. Section 07.
↓    ↓    ↓
System of record Nutshell Leads, contacts, accounts, opportunities, stages, tasks, activities. Authoritative. Native AI stays off.
Transaction E-sign · Payments Bought, not built. The gate drives them and reconciles their webhooks back to the opportunity.
Attention Email · Slack · Calendar Reports, alerts, approvals, escalations and bookings land where your team already looks.
Why the gate matters more than the agents

Consider what happens when the AI is mid-negotiation, a rep is editing the same record, and a webhook fires — all at once. The honest answer is that no amount of prompt engineering solves that, because you cannot make fifteen independent language models agree on a locking protocol. It has to be solved once, deterministically, outside the agents. That is what the gate is, and it is the first thing I would build after the Nutshell adapter.

02 / Surface map

Every Phase 1 human touchpoint, and where it lives

This is the table the Canonical is missing. It takes each Phase 1 capability that a person has to see, touch or approve, and assigns it a venue. Four venues only: Nutshell, a channel you already read, a product you buy, or the Control Plane. Anything that cannot be assigned to one of the first three has to be justified, because every row that lands in the fourth column is custom software somebody has to maintain.

Nutshell Native record, field, list, task or pipeline
Channel Email, Slack/SMS or calendar
Bought A third-party product DMA holds the account for
Control Plane The one purpose-built application
Phase 1 touchpointVenueHow it actually appears
The rep's day
Rep-specific priority list Nutshell A saved filtered list: owner = me, ai_tier = A, ai_next_touch ≤ today, sorted by ai_score. No new screen — the AI writes the fields, Nutshell does the sorting.
Next Best Action per record Nutshell Two custom fields on the record: the instruction, and the one-line reason behind it.
Pre-call brief Nutshell Channel Written into the record as a note when a meeting is booked; also pushed to the rep by email or Slack 30 minutes before the call.
Per-record AI on/off Nutshell A custom field the rep sets: Full · Draft only · Research only · Off. The gate reads it before every action. This is the "don't touch this one, it's mine" control you described.
Automatic tasks & promise detection Nutshell Native Nutshell tasks, created by the gate with a source tag so an AI-created task is distinguishable from a human one.
Live human takeover of a chat Bought Whatever chat product the widget runs on handles operator takeover. We supply the transcript and account context into it; we do not build a chat console.
Management
Daily money report Channel A 7am email. Top 10–25 dormant opportunities, score, reason, recommended action, each deep-linking into the Nutshell record. Same data, no dashboard to remember to open.
Hot-lead alerts Channel Slack (or SMS) with score, company, problem, recommended action and a link. SLA clock starts on delivery.
Neglected-lead escalation & SLA reassignment Channel Nutshell Escalation message to the manager; reassignment written to the Nutshell owner field.
Revenue-leakage detector Channel Control Plane Digest email of gaps — proposal sent with no follow-up, verbal yes with no contract, contract with no payment. The detector's rules and thresholds live in the Control Plane because they are policy, not CRM data.
End-to-end revenue measurement Control Plane Visitor → conversation → qualified → meeting → proposal → client → MRR, with AI-created / AI-influenced / AI-touched separated and a holdout cohort. Nutshell cannot express this; it is the attribution model.
Approvals
Zero-Entry CRM "Approve All" Channel Control Plane The one genuinely contested row. Default: a batched email with signed approve / edit / reject links, so a rep can clear a day's extractions from their phone. Control Plane holds the same queue as a fallback and for anything needing an edit. See the note below.
Staged reactivation batch approval Control Plane The first 100–250 sends need a human to read a sample and release the batch. This is an operator action, not a rep action, and it needs to show deliverability state, suppression counts and the holdout split alongside the drafts.
Escalation (seven triggers) Channel Nutshell Slack message naming the trigger, sent to whoever that trigger maps to, plus the opportunity flipped to an Escalated state with the trigger recorded on the record. Who takes it from there is handled in Nutshell, the way your team already hands work over.
Undoing something the system did Control Plane Restoring a prior value across possibly many records is not something a CRM UI is built for, and Nutshell does not retain prior values at all. The activity log and the restore live here. What is and is not reversible is set out in section 07.
Post-close sampling / audit review Control Plane A rolling sample of autonomously closed deals with the full decision trail. This is the audit layer the Canonical requires; it has nowhere else to live.
Transaction
Proposal delivery Bought Generated from your template and sent through the e-sign or proposal product, so the prospect gets a document from a system they recognise.
Contract signature Bought E-sign provider's own signing experience. We never build a signing page.
Deposit / payment Bought Processor-hosted checkout link. Card handling, retries, receipts and disputes are the processor's problem, by design.
Transaction state & failure recovery Control Plane Three external systems each know one third of the truth. Something has to join them and own the failure cases — signed but unpaid, abandoned checkout, refund request. That join is the Control Plane's transaction console.
Pricing guardrails & negotiation limits Control Plane Discount floors, approved payment terms, scope boundaries, deal-size ceiling, approved-precedent table. Versioned, effective-dated, human-edited. This is policy data and it does not belong in a CRM.
Kill switch & autonomy switchboard Control Plane Global stop, per-engine stop, per-capability autonomy level, per-segment pause. Per-record control stays in Nutshell where the rep is.
Contradiction in the Canonical — "Approve All"

Section 20 makes one-click batch approval a non-negotiable Phase 1 item, and its own v3 amendment says "no custom UI extension — use approved API/native Nutshell writes." Those cannot both hold unless Nutshell supports a bulk-edit-from-list-view flow we can drive, which I have not been able to confirm without the live account. There are three honest answers and Discovery picks one: an email approval digest with signed links (my default — zero new UI, works on a phone, and it is how the rep already handles their day); Nutshell bulk edit on a staging field, if the account supports it; or one screen in the Control Plane. I would rather flag this now than ship you a mock of a screen your own document forbids.

03 / Inside Nutshell

What a rep sees on Monday morning

The Canonical's rule that Nutshell stays the single system of record is the right call, and it has a concrete design consequence: everything the AI knows about a record has to be expressible as Nutshell fields. If it cannot be written to a field, a task, a note or a stage, a rep will never see it, and the system becomes a second place to look — which is exactly the failure mode you are trying to avoid.

So the first real design artefact of this project is not a screen. It is a field schema.

Nutshell  ·  Opportunity  ·  Ridgeline Orthodontics — SEO + Local
Ridgeline Orthodontics Stage: Qualified  ·  Owner: Josh  ·  Value: $6,500/mo  ·  Last human activity: 14 Feb 2024
AI: Full autonomy
ai_score87 / 100 · re-scored 2h ago
ai_tierA
ai_next_actionEMAIL
ai_next_touch29 Sep 2026
ai_next_action_reasonLost Mar 2024 to a 24-month incumbent contract noted in the call log; that term ends this quarter. Two new locations added since. Contact verified, still Practice Manager.
ai_lost_reason_normIncumbent contract
ai_contact_statusVerified · 12 Sep 2026
ai_consent_basisPrior inbound enquiry
ai_stateQueued — reactivation
ai_last_actionEnrichment refresh · 27 Sep 2026 06:12
ai_run_idrun_8f2c41 — opens the activity log
AI autonomy for this record Full · Draft only · Research only · Off — the gate reads this before every single action
Change

Sketch. Field names are placeholders; the real schema is agreed in Discovery against your live Nutshell custom-field capabilities.

The field schema, and why each one exists

FieldTypeWritten byWhat depends on it
ai_scoreNumber 0–100Scoring pipelineTiering, priority lists, hot-lead threshold, routing, SLA class
ai_tierA / B / C / SuppressDeterministic thresholdsWhich outreach cadence a record is eligible for
ai_next_actionEnum, 7 valuesNBA agent → gateRep priority list, daily money report
ai_next_action_reasonTextNBA agentThe Canonical's explainability rule — every recommendation states its evidence
ai_next_touchDateCadence engineFrequency caps, list filters, contract-expiry timing
ai_autonomyFull / Draft / Research / OffHumanRead by the gate before every action. Human always wins.
ai_stateIdle / Queued / Awaiting reply / Escalated / SuppressedGatePrevents two engines working the same record; drives the escalation view
ai_escalation_triggerEnum, the 7 triggersGateAn escalation names the trigger that caused it, rather than arriving unexplained
ai_lost_reason_normEnumCleanup passLost-deal resurrection, deal autopsy, scoring model
ai_contact_statusVerified / Unverified / Departed / BouncedVerification serviceSuppression — no send without a verified path
ai_consent_basisEnumAudit passWhether this record is contactable at all, per jurisdiction
ai_run_idTextGateTraceability — any AI-written value on the record opens its own activity log entry

The approved list of permissions by field

A field schema is only half the answer. The other half is a list DMA gives me: for each field, what this system is permitted to do to it. "The AI writes to Nutshell" is far too coarse a statement to build against, and I would rather be handed the boundary than infer it. It is one short document, it is yours rather than mine, and once it is signed off the gate simply enforces it. Each field carries three properties:

  • May the system write it at all — deal value, close date and owner are candidates for people-only, and saying so up front is cheaper than discovering it after a bad write.
  • At what autonomy — write freely, write only with approval, or propose and never write.
  • What happens on collision — whether a human value is protected, and for how long.

That last property is the one that matters most in practice, and it is the mechanism behind the human-priority rule in section 07. When a person edits a field the system also writes, that field becomes theirs. The system stops writing it, keeps computing its own value, and shows the two side by side rather than overwriting. It is the honest behaviour: a rep who corrects the AI once should not have to correct it again on Thursday.

The obvious version of this leaves the lock in place forever, and over a 30,000-record database that quietly strangles the system — a year on, a third of your records are frozen by corrections nobody remembers making. So the lock has a life: it holds until the account's next substantive human touch, or a set number of days, whichever comes first. When it lapses, the system does not resume writing. It proposes, with the human value shown next to its own and the date the lock was set. The person keeps the last word; they just do not have to keep saying it forever.

The system gets its own name in Nutshell

Every write lands as a distinct DMA AI user rather than under a rep's credentials. That buys three things: a rep looking at a record can tell at a glance what a person did and what the system did; the Control Plane's activity log can be reconciled against Nutshell's own timeline; and the echo-suppression in section 07 has something reliable to key on.

Assumptions I have not been able to test

Everything on this page assumes the live Nutshell account supports: custom fields of these types on the objects we need; API writes to them at the volume of a 30,000-record backfill plus continuous write-back; a change-notification path (webhooks, or polling we can live with); and a way to distinguish our own writes from human ones so we do not trigger ourselves in a loop. I have not verified any of this against your account. A read-only credential and two days settles it, and it is the single highest-severity unknown in the project — a polling-only or field-restricted API adds a state-reconciliation layer that no estimate from any candidate currently includes.

04 / Email, Slack, calendar

The surfaces you already check

Three of the Canonical's Phase 1 "non-negotiables" — the daily money report, hot-lead alerts and one-click approval — read like dashboard requirements. They are not. A dashboard is a place you have to remember to go; the whole point of the daily money report is that leadership sees it whether or not they were thinking about it. These belong in the inbox and in Slack, and building them there is both faster and more likely to actually get used.

Email  ·  07:00 daily
Daily Money Report — 29 Sep DMA Revenue System → jeev@, josh@
1  Ridgeline Orthodontics  87 Incumbent contract ends this quarter · 2 new locations · $6.5K/mo prior scope
Open in Nutshell
2  Kettle & Vine Group  84 Former client, left over reporting cadence · new CMO started Aug
Open in Nutshell
3  Northgate Dental Partners  81 Proposal Feb 2024, never followed up · site rebuilt, rankings down 40%
Open in Nutshell

+ 17 more · yesterday: 4 replies, 2 meetings booked, 1 escalation

Illustrative records. Every row deep-links to the Nutshell opportunity — the report is a pointer, never a second system.

Slack  ·  #dma-revenue
AI
DMA Revenue System 09:41
Hot lead — score 91 · Tamsin Oyelaran, Marketing Director, Crestwell Retail Group
Asked about multi-location PPC for 34 stores. Budget signal: "north of 20 a month." Wants to move before Q4.
Recommended: call today. SLA clock: 60 min.
Open opportunityClaim

!
DMA Revenue System 11:08
Escalation → Josh · Trigger 7: no approved precedent
Harbourline Medical asked for quarterly-in-advance terms on a 12-month retainer. No approved precedent for that payment-term / term-length combination. Deal held at Awaiting decision; nothing sent.
Review & approveAdd as precedent

The escalation names its trigger. "Add as precedent" is the mechanism that makes the system get less needy over time — see section 09.

Where an escalation is sent

The Canonical routes all seven triggers to Josh, and reading it closely that is a heavier load than it first appears: Josh also takes every trust-building video call, by a standing rule set outside the seven triggers. One person as the sole destination for every exception is a bottleneck, and the failure is silent — the escalation fires, that person is mid-call, and the deal sits.

The fix is small and it is a mapping, not a process. Each of the seven triggers points at one or more destinations, and that mapping is a setting in the Control Plane that DMA edits directly — no deploy, no ticket to me. That is the entire extent of what the system needs to know.

Deliberately, it knows nothing more. It does not track who is on holiday, run a timer, chase an acknowledgement or re-route to a backup, because Nutshell and Slack already do the human half of this better than a rule I would write in advance: the opportunity is flipped to Escalated and assigned, and from there people tag each other, watch a colleague's queue while they are out, and hand things on. Encoding coverage rules into this system would freeze an arrangement your team changes every week, and it would quietly become another place to keep up to date.

What the system owes you instead is visibility, which costs nothing to provide: escalated opportunities are a state in Nutshell, so "what is escalated and how long has it been sitting there" is a saved list rather than a feature. If DMA later wants a clock on that, it is a Nutshell automation over a field that already exists — not something to build here now.

One route does need to be populated at all times, and it is the only hard requirement: the prospect who asks for a human. That trigger has no confidence threshold and no deferral, so its mapping cannot be empty. Everything else can be left to the team.

Approving from an email, safely

If the default approval surface is a link in an email, that link is a credential and has to be treated as one. Each is single-use, short-lived, bound to one named approver, and bound to the exact version of the thing they were shown — so if a draft changes between the digest going out and the click landing, the link no longer approves anything and the item returns to the queue. Approving is the only thing these links can do; they are not a way into the Control Plane. A forwarded email should never be able to release a batch of outreach.

Calendar

Real booking is a Phase 1 item and it is a solved problem: the agent reads availability and writes the event through the calendar API, against the routing rules for service lane and deal size. The one design decision worth making early is that the agent books into a real calendar, not a booking-link handoff — the Canonical is right that a booking link is a drop-off point, and once the agent has the prospect's attention it should not let go of it.

05 / Bought, not built

The services, and who should own the accounts

Every one of these is a category where building is strictly worse than buying. My recommendation across the board is that DMA holds every account directly — you own the relationships, you see the spend, and if you ever change development partners nothing has to be untangled. I would quote integration effort only, and exclude all pass-through fees and model spend from any price. That also keeps your comparison between candidates honest, since a bundled number hides which is which.

NeedShape of the answerIntegration effortNotes
E-signaturePandaDoc / DocuSign classMediumNeeds template + merge fields + completion webhook. Choice partly depends on whether you want proposal and contract in one product.
PaymentsStripe classMediumHosted checkout, not custom. Retry, dunning and dispute handling come with the product and should not be reimplemented.
ESP + sending domainDedicated sending domain, SPF/DKIM/DMARC, warm-upMediumMust be configured and warmed before any reactivation send. This is on the critical path for the first revenue milestone.
Email verificationVerification + finder serviceLowGates every send. Directly determines how much of the 30K is reachable at all.
Firmographic enrichmentClearbit classLow–MediumFeeds scoring and personalisation. Per-record cost across 30K is a real budget line.
SEO / site analysisLikely covered by DMA's existing toolingLowFeeds the Mini-Consultant. Worth confirming what you already pay for before buying anything new.
Website chat widgetExtend existing, or embed newUnknownDepends entirely on your website platform, which I do not yet know. Also determines how live human takeover works.
CalendarGoogle / Microsoft APILowDirect API, not a booking-link product.
AlertingSlack, optionally SMS for hot leadsLow—
Model providerProvider-flexible; routing by task—Cheap models for bulk extraction and classification, strong models for customer-facing copy and negotiation. Costed below.

What the model spend actually looks like

These are estimates, and I want to be exact about what kind. The unit prices are published list prices. The token volumes per task are my engineering estimates for work of this shape, not measurements — I have not run your data through anything. The monthly volumes are assumptions about DMA, and you should correct them. Everything below can be recalculated once any of those three change, which is the point of showing the arithmetic rather than a single number.

Two model tiers, chosen per task. A cheap, fast model does the bulk reading — 30,000 records, transcripts, inbound replies, classification. A stronger model writes anything a prospect will read. Mixing them is not a cost trick; sending a reactivation email through the cheap model to save a third of a cent is a bad trade against the reply rate.

Unit of workTierEstimated cost each
Read and score one dormant recordBulk$0.004 – $0.020
Write the "why now" line that travels with a reactivation candidateWriting$0.008 – $0.014
Draft one personalised reactivation emailWriting$0.010 – $0.025
Classify an inbound reply and draft the answerBoth$0.012 – $0.020
One whole website conversationWriting$0.060 – $0.150
One live site review for the Mini-ConsultantWriting$0.030 – $0.060
One call transcript into CRM fields, tasks and a follow-up draftBoth$0.020 – $0.035
Re-score a record after it changes in NutshellBulk$0.003 – $0.008
Operating tempoAssumed volumeEstimated model spend
Reactivation at a low cadence800 drafts · 100 conversations · 60 site reviews · 50 calls · 1,000 re-scores$40 – $90
Reactivation at full cadence2,000 drafts · 300 conversations · 200 site reviews · 150 calls · 3,000 re-scores$90 – $200
Full cadence over a heavy inbound month5,000 drafts · 1,000 conversations · 500 site reviews · 300 calls · 6,000 re-scores$230 – $520
One-timeBasisEstimate
First full pass over 30,000 records30,000 × the per-record range, plus a written reason for each record that clears the tier B threshold. How many that is, is one of the things the first pass tells you — the estimate spans a quarter to a half of the database$200 – $750
The same pass, submitted as batch workNon-urgent bulk work runs at half price, and the audit is the definition of non-urgent$110 – $400
Prompt tuning and regression runs before go-liveRepeated evaluation passes over a fixed set of DMA examples$50 – $250
The model is the cheap part — and that is the finding

Model spend for a system of this shape lands in the low hundreds of dollars a month. The data vendors in the table above do not. Firmographic enrichment across 30,000 records is a one-time charge in the hundreds to low thousands before any ongoing refresh; email verification adds a similar one-time hit; an ESP with a dedicated domain, an e-sign seat and SEO tooling are each a monthly line in the tens to hundreds; payment processing is a percentage of revenue. Realistically the third-party stack costs several times what the AI does, and it is the line that actually needs a decision from DMA. A quoted running cost that counts only model tokens is measuring the smallest item on the bill.

What stops it running away

None of these depend on a model behaving. A hard monthly ceiling that refuses further calls and falls the website chat back to a plain capture form. A separate daily ceiling per subsystem, so a runaway bulk job cannot starve the live chat. A per-conversation cap that hands to a human rather than spending indefinitely on one visitor. Rate limits per visitor and per conversation length. And the expensive reads are keyed on their inputs rather than on the clock — the same content fingerprint the adapter uses in section 07 decides whether anything needs looking at twice, so the bill tracks how much your data actually moved rather than how often the system woke up. That is the whole reason the recurring number is a fraction of the one-time one: after the first pass, you are paying for change, not for volume.

06 / The Control Plane

The only thing here that has to be purpose-built

Everything in the previous three sections was Nutshell, a channel, or a product you buy. What is left is the Control Plane, and the test I applied to every item in it is simple: could this live in Nutshell, in an inbox, or in a product we can buy? If yes, it is not in here. What survives that test is eight things, and none of them are sales work.

This matters for the Canonical's "no custom approval UI" rule. The rule is about not making reps learn a second CRM, and this design honours it — a rep can do their entire job without ever opening the Control Plane. What the rule cannot mean is that a system with autonomous pricing, contracts and payments has no operator console at all, because then nobody can turn it off, nobody can change a discount floor, and nobody can answer "why did it do that?" six weeks later.

01 · PolicyGuardrail registryPricing floors by service and tier, the discount ladder, approved payment terms, scope boundaries, the deal-size ceiling, and the approved-precedent table. Versioned and effective-dated, so a decision made in March can be replayed against March's rules.
02 · SafetyAutonomy switchboard & kill switchGlobal stop; per-engine stop; an autonomy level per capability; pause by segment or cohort. One button that halts everything outbound, reachable in under five seconds.
03 · WorkAction queueEvery pending intent with the evidence behind it, the guardrail checks it passed, and what it will do. Batch release for staged sends. Where a human interrupts the machine before it acts, rather than after.
04 · AuditActivity log — the audit viewAgent, prompt version, model, input snapshot, structured output, cost, committed action, outcome. This is what makes post-close sampling possible and what answers "reproduce exactly what happened" a week later.
05 · MoneyTransaction consoleProposal, signature and payment state per deal, joined from three external systems, with the failure cases made explicit and actionable. Section 10.
06 · HealthRun health & spendNutshell sync lag, webhook or poll status, queue depth, reconciliation drift, bounce and complaint rates, SLA clocks, model spend against its ceiling. The dashboard I look at, not one your reps do.
07 · RecoveryActivity log — the restore viewThe same entries, read for recovery rather than audit: each carries the prior value of whatever it changed, so a field can be put back by re-applying it through the gate — one field, one run, or a named batch. Section 07 is honest about what can and cannot be taken back.
08 · KnowledgePlaybook, prompts and proof libraryThe sales playbook, objection library, competitive positioning and approved proof points, each versioned, with a review step before any change takes effect. This is the Canonical's "human-reviewed model changes" rule made real: the system never silently rewrites its own sales strategy, because a change to the strategy is a reviewed commit with an author and a date.

Action queue — the operator's screen

Control Plane  ·  Action queue  ·  41 pending
Reactivation batch 03  Awaiting release 220 recipients · Tier A, incumbent-contract cohort · 31 suppressed (unverified 19, opted out 7, active client 5) · 48 held as holdout Domain health: good · 7-day bounce 0.4% · complaint 0.01%
Read 10 draftsRelease batch
Proposal — Crestwell Retail Group  All checks passed $21,400/mo · 12-month term · standard payment terms · discount rung 2 of 4 · precedent: 6 matching approved deals Claims check: 4 figures cited, all resolved to account record or approved proof library
ViewSend autonomously
Proposal — Harbourline Medical  Held · trigger 7 Quarterly-in-advance on a 12-month term. No approved precedent for this combination. Nothing has been sent.
RejectApprove & add precedent
17 Zero-Entry extractions  Nutshell writes From 6 Zoom transcripts · budget, timeline, services, decision maker, objections · also queued to reps as an approval email
ReviewApprove all

Sketch of placement and content. Note what is not here: no lead list, no pipeline view, no contact records. Those are Nutshell's, and duplicating them is how you end up with two systems of record.

07 / The Commit Gate

What happens when everything moves at once

This is the hardest part of the project, and the part most likely to be quietly missing from a system that demos well. An agent is drafting a reply, a rep is editing the same record, a prospect's email arrives and Nutshell fires a webhook — all inside the same few seconds. Who wins, what gets thrown away, and what must never be allowed to happen twice? Here is the mechanism, in full.

No agent ever writes directly. An agent produces an intent: a proposed action, the snapshot of state it was computed from, the evidence behind it, and an idempotency key. The intent goes on a queue. The gate is the only code in the system that touches Nutshell, sends mail, or calls the payment and e-sign providers, and it runs these steps in order every time.

1
Re-read live state Fetch the current Nutshell record immediately before acting. Not the copy the agent worked from — the live one.
2
Compare against the snapshot, by materiality Each intent type declares which fields matter to it. A phone-number edit does not cancel a queued outreach email. A stage change to Closed Lost, a new human note, an owner change or an autonomy change does. → Material change: discard the intent and re-plan from current state. Never retry the stale one.
3
Human-priority check Any human activity on the record inside the lookback window invalidates pending machine intents against it. If a rep left a note saying they had spoken to the prospect, the queued "just checking in" email dies and the correct response becomes a task, not a send. This is also where field ownership from section 03 is enforced: a field a person has taken over is not written, only proposed. → Human acted: discard, re-plan, and in most cases the re-plan is "do nothing, create a task."
4
Autonomy and suppression check Global kill switch, engine switch, capability autonomy level, the record's own ai_autonomy field, suppression list, unsubscribe, frequency cap, quiet hours, active-client check, open support or billing issue. Circuit breakers sit here too: bounce and complaint rates crossing their thresholds pause sending automatically rather than waiting for someone to notice on a dashboard, because by the time a deliverability problem is visible in a weekly review the sending domain is already damaged. → Blocked: log the reason, leave the record untouched.
5
Guardrail evaluation — deterministic Price, discount rung, payment terms, scope, term length and deal size are checked in code against the versioned guardrail registry. The language model proposed the numbers; it does not get to approve them. → Out of bounds: escalate. Never silently clamp to the nearest legal value.
6
Escalation triggers — deterministic All seven Control Plane triggers evaluated in code. Any hit routes to Josh with the trigger named, and the opportunity moves to Escalated.
7
Claims check on anything customer-facing Every figure, ranking, result or guarantee in outbound copy has to resolve to a fact on the account record or an entry in the approved proof library. Unresolved claim, blocked send. The Canonical's "no invented claims" rule is a validator, not a line in a prompt.
8
Commit once, under a per-record lease A short-lived lease on the record means two agents cannot act on the same opportunity concurrently. Every write goes through a single Nutshell adapter that rate-limits, backs off on failure, captures the value it is about to replace, and stamps the write as ours. Exactly-once is harder than it sounds against a third-party API and is handled below.
9
Log, then hold Agent, prompt version, model, input snapshot, output, checks run, decision, cost, and later the outcome. Written with the same ai_run_id stamped on the Nutshell record, so any field a rep is looking at can be traced back to the reasoning that produced it. Reversible classes of action then sit in a short hold window before they take effect, which is the only free chance anyone gets to stop them.

Writing exactly once, against an API that may not help

The textbook answer is an idempotency key: attach a unique token to the write, and the far end guarantees it applies once no matter how many times you send it. If Nutshell supports that, we use it and this problem is over.

It may not, and most CRM APIs do not. The failure that matters is narrow and specific: we send a write, the connection dies before the response arrives, and we genuinely do not know whether it landed. Retrying blindly gives your prospect two identical notes, or two tasks, or a duplicated activity. So the adapter carries its own answer. Every created object gets a deterministic fingerprint derived from its content and the run that produced it, written into the object itself. Before any retry, the adapter looks for that fingerprint. Found means the first attempt succeeded and the retry is recorded as a no-op. Not found means it genuinely failed and the write proceeds. For field updates the problem is easier — if the field already holds the value we meant to write, we are done.

I am flagging this rather than assuming it away because it is exactly the kind of thing that works perfectly against a simulator and then does not survive the real API.

When Nutshell tells us something changed

Change notifications are treated as signals, never as data. A webhook tells us a record is worth looking at; it never tells us what the record now says. The system goes and reads the live record before doing anything with it.

That one rule disposes of a whole family of problems at once. Events arriving out of order stop mattering, because whichever arrives last still triggers a fresh read of the same current truth. A replayed or duplicated event costs one redundant read and is otherwise inert, and duplicates are dropped on event identity anyway. And a notification whose payload is stale — which is normal under load — can never roll a record backwards, because we never trusted the payload.

Our own writes come back as notifications too. The activity log knows what the system just wrote, so an echo of our own change is recognised and consumed silently rather than waking an agent. Without that, the system responds to itself, and a small loop becomes a large one very quickly.

The nightly argument with Nutshell

Everything above assumes we are told when things change. Assume we are not. Notifications get dropped, a polling window gets missed, someone does a bulk import, an integration nobody mentioned writes to the same fields. Any working mirror of a system you do not control needs a reconciliation loop, and it needs to do something when it finds a problem.

So a full pass compares the mirror against Nutshell nightly, and a continuously running sample checks a small slice throughout the day — because a bad deploy at 9am should not have until midnight to do damage. Discrepancies are sorted into two kinds, which matter differently. A value we wrote that no longer matches means a person changed it and we missed the notification: the mirror is corrected and that field becomes theirs, exactly as if we had seen the edit. A value neither side can account for is a genuine integrity problem and is the one that should make somebody nervous.

Drift is measured per field class rather than per record, because the thresholds are not the same: a percent of contact-detail churn overnight is ordinary, while the same drift in deal stages or values is not. When a class passes its threshold, writes for that class pause automatically and the queue holds rather than drains. Nothing is dropped, nothing is lost, and someone has to look before it resumes. The alternative — a system that keeps confidently writing into data it has quietly lost track of — is how a CRM integration turns into a data-recovery exercise.

What can actually be undone

"Undo" is the first thing anyone asks for and the word covers two very different things, so it is worth separating them plainly.

Writes into Nutshell can be put back — but not by Nutshell, and I checked that rather than assuming it. Nutshell's own Audit Log is an Enterprise-plan admin screen covering logins, bulk edits and list exports; it cannot be exported and has no API behind it. Its change feed is more promising but stops short of what an undo needs — Nutshell's own guidance is to treat a change notification as a prompt that an entity moved, not as a description of what moved, and the sample payload ships an empty changes array with only the entity's current state. Nothing in Nutshell can tell you what a field held before it was overwritten. So prior values have to be retained on our side, and the only real design choice is how much apparatus that justifies.

My answer is: almost none, because the Canonical already requires the expensive part. Every AI decision has to be traceable and costed, which means the Control Plane keeps an activity log regardless — one entry per action, with the agent, the rule that allowed it, the cost and the outcome. The gate already pulls the live record immediately before it writes (step 1), so the outgoing value is in hand at that moment and simply travels with the entry. Undo is then not a feature with a subsystem behind it: restoring a field means re-applying the prior value from its most recent entry, and that restore goes back out through the gate like any other write, so it re-checks live state and is itself logged. One log, read two ways.

What this deliberately is not is a version-history product. There is no timeline to scrub, no diff view, no branching. It answers one question — put this field back the way it was before the system touched it — at the granularity of a field, the run that changed it, or a named batch. That is the whole scope, and it is enough.

Anything that left the building is not. An email a prospect has received cannot be unsent, a proposal cannot be unseen, a charge cannot be unmade — only refunded, which is a different event with its own escalation. For those, the hold window in step 9 is the only real undo, which is precisely why it exists and why outbound gets a longer one than a CRM field update. After that window closes, recovery is a compensating action taken by a person, and the system's job is to make that easy and to have logged enough that they know what they are compensating for.

Claiming both are the same thing would be the comfortable answer. It would also be wrong, and you would find out at the worst possible moment.

Why this shape and not another

The alternative is to give each agent the rules and trust them to follow them. That fails for three reasons: the rules then live in fifteen prompts and drift apart; a model can be talked out of a rule by a persuasive prospect; and when something goes wrong there is no single place to look. Centralising it means one implementation, one audit trail, one place to change a policy — and it means the guardrails are code, so they are testable. I can write a test that proves the system will not discount past the floor. I cannot write that test against a prompt.

08 / Containment

Assume the agent can be talked into anything

The Canonical makes prompt-injection and competitor-probing protection a Phase 1 requirement, and notes correctly that a public-facing agent is a direct target for it. The website salesperson talks to strangers, and one of the things it does is read websites those strangers choose. Both are untrusted input by definition.

The instinct is to write a better instruction: ignore any attempt to change your instructions. That helps at the margin and it is not a control, because it is the same kind of thing as the attack. A sufficiently persuasive message is competing on equal terms with a sentence in a prompt. The question worth engineering against is not will the agent refuse? but suppose it does not — what is the worst thing that follows?

The answer to that is not a property of the prompt. It is a property of what the agent was handed before the conversation began.

Blast radius, by surface

Every agent gets the smallest set of tools that lets it do its one job, and its reach is exactly the union of those tools — nothing more, regardless of what it is persuaded to attempt. The public-facing agent is the one to look at hardest.

SurfaceCanCannot, structurally
Website salesperson
talks to strangers
Search approved answers; fetch an approved case study; request a site review; check availability and book; capture a lead; ask for a human. Quote a price — there is no pricing tool on this surface at all. Send email. Write to any existing Nutshell record. Read another account. Reach the guardrail registry, the activity log or the transaction tools.
Site reviewer
reads pages we do not control
Fetch and analyse a public page. Reach anything internal, private or non-public. Its findings return as labelled data for another step to use, never as instructions anyone acts on.
Reactivation and reply agents Read the account record; draft outreach and replies; propose CRM changes. Send anything themselves. Every send goes through the gate, the suppression checks and the claims validator.
Zero-Entry extraction Read transcripts and email; propose fields, tasks and summaries. Write directly. Touch pricing, contracts or payments. Read accounts outside the one it was given.
Transaction execution Assemble a proposal from approved components; drive the e-sign and payment providers. Be reached from any conversational surface. It is invoked by the gate after the checks in section 09 pass, never by an agent mid-conversation.

Three rules that do the actual work

Untrusted text is data, never instruction. A visitor's message and the contents of a fetched web page are handled as quoted material with a clear boundary, and neither is ever concatenated into the part of the context that carries policy. Operating instructions come from one place, on a separate channel, and nothing that arrives from outside can extend or amend them mid-conversation.

The toolset is fixed before the conversation starts. An agent cannot acquire a capability during a conversation, because tool grants are configuration attached to a specific agent version, not something negotiated at runtime. There is no sequence of messages that gets the website chat an email tool.

Least privilege goes all the way down. Each capability group holds its own credentials against the database and against Nutshell, scoped to what it actually needs. A flaw in the component that handles public chat does not become read access to your 30,000 records, because that component was never able to read them.

The other half: what it must not say

Containment limits what the agent can do. Competitor probing is about what it can be induced to reveal — pricing logic, the playbook, how leads are scored, what it was told about a competitor. Two things hold there. The scoring model, the guardrail registry and the internal playbook are not in the public agent's context to begin with, so there is nothing to extract; it can reach approved public content and nothing else. And the claims validator in step 7 of the gate runs on the way out as well as the way in: anything the agent produces that asserts a figure, a ranking, a result or a guarantee has to resolve to approved source material, which catches both the invented claim and the leaked internal one. Output validation is the control that does not care how clever the input was.

09 / Code vs. LLM

Which decisions a model is allowed to make

The sharpest form of this question is how "no approved precedent → Josh" gets implemented without simply asking a language model whether something feels unprecedented. That one is answered on its own below.

DecisionDecided byMechanism
Whether a record may be contacted at allCodeConsent basis, suppression list, unsubscribe, frequency cap, quiet hours, bounce state, active-client check. No model input.
Tier boundaries and eligibilityCodeFixed thresholds over the score. Changing a threshold is a versioned config change with an author.
The opportunity scoreHybridThe model extracts features from free-text notes, emails and transcripts; the score is a formula over those features with versioned, human-approved weights. The number is reproducible; the reading of the prose is not, which is exactly the right split.
The price itselfCodeThe model never produces a number. It selects services and quantities; code prices them from the approved rate card. See below.
DiscountCodeA lookup against the approved ladder. The model may select a rung it can justify; it cannot invent one, and anything past the floor escalates rather than being clamped to it.
Payment terms, contract template, term lengthCodeAllow-list. Anything not on it is by definition non-standard, which is escalation trigger 2.
Deal-size ceilingCodeNumeric comparison.
Whether to send a proposalCodeA gate over a model-drafted artefact. All seven triggers plus the claims check, all evaluated in code.
Whether there is approved precedentCodeA table lookup. See below — this is the one worth spelling out.
Refund or cancellationCodeAlways escalates. No exceptions, no confidence threshold.
Payment retry and dunningCodeThe payment processor's own logic, driven by a state machine. Explicitly not an agent — an agent has no business deciding when to re-charge a card.
Which sales methodology to runModel, boundedSelection from a fixed set of five, driven by signals the system already produces. The selection is logged so it can be reviewed against outcomes.
Tone, phrasing, objection responsesModelBounded by the playbook, the objection library, and the claims validator on output.
Summarising, extracting, classifyingModelStructured output against a schema. A response that fails schema validation is a failed run, not a partial write.
Merging duplicate recordsHumanDetection is automated, merging is not. Merges are destructive and hard to unwind — your own Canonical corrected this one, and it is right.

The system does not compose prices

There is a weak version of a pricing guardrail and a strong one, and the difference is worth being explicit about because they sound alike.

The weak version lets the model write a number and then checks whether that number is allowed. It mostly works. It also means a figure exists, inside the system, that no rule produced — and the entire safety of the arrangement now rests on the validator being complete. Every gap in it is a price nobody approved.

The strong version removes the model from the job. It decides what the prospect needs: these services, this scope, this term. Code turns that into money, line by line, from the rate card, applying the discount ladder. The number is computed, reproducible, and attributable to a specific version of a document DMA wrote.

The engineering that enforces this is unglamorous and effective: the model's output schema has no price field. It cannot emit what it has no slot to emit. And on the public website surface there is no pricing capability at all — an inbound visitor asking "what would this cost?" gets a range only if approved public ranges exist for that service, delivered by code, or gets a conversation about scope and a booked call. Nobody has to trust the agent to be careful with your pricing, because it was never handed it.

"No approved precedent" as a lookup, not a judgement

Asking a model "does this seem unprecedented?" is unanswerable, because the model has no reliable memory of what DMA has approved before and every incentive to be agreeable. So the system does not ask it.

Instead, every autonomous commercial action has to resolve to a precedent record. A deterministic classifier reduces the action to a key made of a fixed set of attributes — service mix, deal-size band, contract term, payment terms, contract template version, discount rung, and any non-standard scope flags. The model's only job is to fill those slots from the conversation; it does not get a vote on what the slots mean. The gate then looks the key up in the approved-precedent table, which is seeded in Discovery from your historical closed-won deals and your written pricing policy.

  • Key found — this shape of deal has been approved before. Proceed autonomously.
  • Key not found — escalate to Josh, naming the attribute that had no precedent. Nothing is sent.
  • Josh approves — the key is added to the table with the approver's name and the date on it. The next deal of that shape proceeds without an escalation.

That last line is the important one. The system starts conservative and becomes autonomous as a direct function of decisions a human actually made, rather than as a function of how confident a model happens to sound. And because the table is data, you can look at it, argue with it, and take entries out.

Who is allowed to change the rules

All of the above is only as good as the change control around it. A discount floor that any one person can quietly edit is a suggestion. So every load-bearing setting — the guardrail registry, the precedent table, the escalation routing, the field permissions, and the agents' own instructions — is treated as code rather than as configuration: an edit is a proposal that sits inert until a second named person signs it off, and it carries that person's name for as long as it is in force. Authorship and authorisation are never the same signature, mine included.

Two properties follow that matter more than the rule itself. Settings are effective-dated rather than merely current, so a decision taken in March is replayed against the floors and terms that were in force in March — without that, your audit trail explains what happened using rules that did not yet exist. And because each signed-off state is retained, withdrawing a bad change is a one-line operation with a name attached, instead of a conversation about what the number used to be.

This is also the Canonical's human-reviewed model changes requirement, which asks that the system never silently rewrites its own sales strategy. Made concrete, that is exactly this: a change to the strategy is a reviewed commit with an author, an approver and a date.

Contradiction in the Canonical — triggers 1 and 7

The Control Plane says the AI hands a deal to Josh when one of seven conditions is met. Trigger 7 is "no approved precedent." Trigger 1 is "deal exceeds the largest previously-approved deal size" — and then says "Not blocked — the first deal at a new tier proceeds and gets sampled review after close." A deal above every previously approved size has no approved precedent by definition, so trigger 7 says stop and trigger 1 says go, on the same deal. This is not a nitpick: it decides whether your single largest-ever deal gets sent autonomously. My recommendation is that trigger 1 holds and trigger 7 yields — a new size tier is exactly the moment a human should look — but that is DMA's commercial call, not mine, and it needs making before the transaction layer ships.

10 / Transaction path

Conversation to collected money, including when it breaks

The four v3 additions — methodology orchestration, autonomous proposal sending, e-sign execution and deposit capture — are the newest and least specified part of the Canonical, and they are where the real risk sits. The happy path is straightforward. The value is in the branches.

QualifiedScope & price assembledFrom the Account Digital Twin: conversation, history, prior proposals, service fit. Pre-filled, not re-typed.
GateGuardrails + precedentDeterministic. Pass → proceed. Fail → Josh, with the reason named.
SentProposal via e-sign productYour template, your branding, a system prospects recognise. Opportunity stage and timestamp written to Nutshell.
↓
NegotiationWithin limits, with memoryConcessions already offered or refused are tracked, so the system never re-offers a rejected discount or undercuts an accepted term.
SignedE-sign completion webhookExecuted contract filed; opportunity advanced; deposit request issued.
PaidProcessor webhook → Closed WonOnly a payment event that matches the expected amount closes the deal. A signature alone does not, and neither does a payment for the wrong figure.

Four rules specific to money

The gate in section 07 governs everything. These four apply only to the transaction path, because the cost of getting them wrong is different in kind.

  • What was approved is what goes out. The approved document is fingerprinted at the moment of approval and checked again at the moment of sending. If anything about it changed in between — a regenerated draft, a rate card edited underneath it — the send stops and returns for approval. An approval is for a specific artefact, not a general blessing.
  • Whoever requested it does not release it. On anything that binds DMA or moves money, the person who asked and the person who approved are two different people. This is the same separation as the rule changes in section 09, applied where it matters most.
  • The provider tells us, and we check anyway. Signature and payment callbacks are verified as genuinely from the provider before anything acts on them, and a reconciliation job asks each provider for the current state of anything that has been open too long. A callback that never arrives is a normal event, not an exception — designing as though it cannot happen is how a paid deal sits in Negotiation for a fortnight.
  • Card data never touches this system. Payment runs through the processor's own hosted checkout. This is not primarily a cost decision; it keeps DMA's compliance surface where it belongs.

The branches that actually matter

What goes wrongWhat the system doesWhere a human sees it
Signed, payment fails Opportunity moves to a distinct Signed — unpaid state so it never sits misleadingly in Closed Won. The processor runs its own retry and alternative-method flow. The system sends a short, non-accusatory follow-up and tries an alternate method. Transaction console; Slack to the owner. Trigger 3 escalates on any dispute.
Signed, checkout abandoned entirely Timed follow-up sequence with a fresh link, then hard escalation. It does not time out and quietly drop the deal — you have a signed contract and no money, which is a human problem. Slack escalation; leakage digest; console.
Prospect asks for a human Immediate handoff, full context, no further autonomous messages on that opportunity. Trigger 4, no confidence threshold, no exceptions. Slack to Josh; opportunity flagged.
Video call scheduled Standing rule from Jeev, independent of the triggers: Josh takes the call. The agent prepares the prospect, briefs Josh, and resumes afterwards from the transcript. Calendar; pre-call brief.
Refund or cancellation requested Always escalates. The system takes no position and sends nothing. Slack; console.
Rep intervenes mid-flow Human priority. Pending intents on that opportunity are discarded and re-planned from the state the rep left behind. Nothing to see — which is the point.
Dependency inversion in the Canonical

Autonomous proposal sending is Phase 1. Its listed dependency, proposal data pre-fill, is Phase 2, as is proposal first draft — which the document also notes is a duplicate of Section 11's automatic proposal draft, itself Phase 2. So the capability that sends proposals ships a phase ahead of the capabilities that build them. Separately, e-sign execution is Phase 1 and depends on "approved contract templates," but no capability in the document produces those; they are DMA authoring work with no named owner. Both are resolvable — pull the pre-fill and draft items into Phase 1, and name a template owner — but they need resolving before this cluster can be scheduled, let alone priced.

11 / Build sequence

What is demonstrably working, and when

A milestone view at Day 10, 20, 30 and 45, with the precondition for each stated plainly, because a milestone whose precondition has not been met is a date nobody should trust. Days are counted from build day one; Discovery precedes day one and runs five to eight days of its own.

The sequencing principle: the 30,000 records are the fastest path to real revenue, so they come first, and nothing downstream gets turned on until the thing above it is trusted. We are not turning the whole pipeline on at once and hoping.

ByDemonstrably workingRequires
Day 10
Foundation
Nutshell adapter proven against the live account — reads, writes, rate limiting, retry, exactly-once handling, change detection, echo suppression. All 30K mirrored, incremental sync running, and the nightly reconciliation with its drift thresholds already switched on. Activity log with prior-value capture, restore, kill switch and autonomy switchboard live from the first day, not retrofitted. First real data-quality numbers on the 30K: record age, field completeness, lost-reason population, contact validity, duplicate density.

Shown as: a live sync you can watch, a write made and then put back from the activity log, and the first honest numbers behind the reactivation thesis.
Nutshell API credential. Agreed field schema. Nothing from DMA beyond access.
Day 20
Database becomes a work queue
Scoring pipeline over all 30,000: score, tier, Next Best Action and its reason, written back to Nutshell under the approved list of permissions by field. Rep priority list as a saved Nutshell list. Daily money report in production. Hygiene passes run: duplicate detection, dead contacts, lost-reason normalisation, stale deals, orphaned opportunities. The scoring model backtested against deals DMA has already won and lost, so we can say whether it would have ranked your real wins near the top before we ask anyone to trust it on a dormant record.

Shown as: Josh opens Nutshell and has a ranked list with reasons attached; leadership gets the 7am email; and the backtest either supports the model or tells us what to fix.
Day 10. Two or three working sessions on what "good lead" means, because the first scoring pass will disagree with you and that disagreement is the useful part.
Day 30
First real revenue motion
Sending domain warmed and authenticated. Verification and suppression enforced. Personalised reactivation drafts released in a staged first batch of 100–250 with human release. Replies ingested, classified, answered, objections handled, meetings booked — all through the Commit Gate. Hot-lead alerts and the SLA clock live. Holdout cohort in place so lift is measurable rather than assumed.

Shown as: real sends to real people, real replies handled, first meetings on Josh's calendar, measured against a control group.
Day 20. Sending domain. Consent position settled. ESP and verification accounts. A decision on what a "successful" reactivation gate means — a send, a reply rate, a booked meeting, or closed revenue.
Day 45
Inbound engine + transaction path
Website agent live: conversation, Mini-Consultant, objection handling, live re-scoring, booking, hot-lead alert, Nutshell write-back. Zero-Entry CRM's five items running off existing Zoom transcripts. Transaction path complete end to end in staging against an e-sign sandbox and payment test mode — proposal, signature, deposit — with all seven triggers wired and the precedent table seeded, and the failure branches from section 10 deliberately exercised.

Shown as: a full transaction walked end to end on a test deal, and then walked again while we break it on purpose.
Day 30. Website platform access. Pricing policy written. Contract and proposal templates. Approved proof and objection libraries. E-sign and payment accounts.

Day 45 is not "autonomous closing is on"

I want to be plain about this because it is where I expect my answer to differ most from a competing bid. Having the transaction path working is not the same as trusting it with your customers and your bank account. Autonomy gets raised as a deliberate, reversible ramp, per capability, and the ramp is measured in weeks of real deals rather than a date on a plan.

Step 1
Prepare and holdEverything is drafted, checked and queued. A human releases every one. This is where the prompts, the guardrails and the claims validator actually get tuned, against real deals.
Step 2
Autonomous within precedent, sampledActions whose key resolves to an approved precedent proceed. Everything else escalates. A rolling sample is reviewed after close.
Step 3
Autonomous by classWhole classes of deal — by size band, service mix and term — go fully autonomous once their sampled review is clean over a meaningful number of deals. Raised one class at a time, lowered instantly if the sampling turns up anything.
Always
Human override, kill switch, auditAvailable at every step and at every level. Not a checkpoint the AI passes through — a right you retain and can exercise in five seconds.
The part nobody's plan can compress

Building this is the predictable half. The unpredictable half is the iteration: the first scoring pass will rank one of your best-ever clients a C, and the useful question then is whether the model is wrong or whether it has found something about your intake that you want to change. That conversation happens repeatedly, it needs you in the room, and it is the difference between a system that runs and a system that earns. I have built this shape of thing before — most recently against Pipedrive, which we use ourselves — and the build is not what takes the time.

12 / Open items

What the Canonical does not yet resolve

The Canonical's own last page says it does not select vendors, design the schema, or produce a budget and timeline. It is a strong decision document and I have treated it as binding. These are the places where following it precisely produces a question rather than an answer — raised now, because finding them in week three is considerably more expensive.

ItemThe tensionEffect if unresolved
Scope of Phase 1 92 rows are marked Phase 1, of which at least four pairs are duplicates of each other. Is the release all 92, or the eleven non-negotiables plus the four v3 transaction items? Blocks pricing Roughly a fivefold swing in effort. The largest single variable in any bid you receive.
Nutshell API reality Every capability rests on two-way access. Write support for the assumed objects and fields, change notification versus polling, per-write idempotency support, bulk-edit behaviour, and rate limits under a 30K backfill plus continuous write-back are all untested. Highest severity A polling-only or field-restricted API changes the sync design rather than adding to it. Worth saying plainly: this is the one question no prototype can answer. A system demonstrated against a stand-in CRM — ours or anyone's — has proved that the software works against the assumptions its authors made about Nutshell, which is a different claim from working against Nutshell. Two days with a read-only credential settles it, and it should be settled before anybody signs anything.
"No custom UI" vs. four Phase 1 surfaces Batch Approve All, the daily money report, rep priority lists and hot-lead alerts are all Phase 1, and the document forbids a custom approval UI. Section 02 proposes an answer for each; they need confirming against the live account. Decides whether there is a web application in Phase 1 at all. Probably the second largest line item available.
Triggers 1 and 7 A new deal-size tier has no approved precedent, so one trigger says proceed and the other says escalate, on the same deal. Decides how your largest-ever deal is handled. Needs a decision before the transaction layer ships.
Proposal dependency chain Autonomous proposal sending is Phase 1; the pre-fill and draft capabilities it depends on are Phase 2. Either pull two items forward or move the sending item back. Straightforward once someone decides.
Shadow pipeline references v3 reverses the shadow pipeline to NO, but estimated latent pipeline value and promote when evidence becomes strong both still list it as their dependency. Low severity. The amendment explains the intent; the dependency rows were not updated.
Transcript extraction phasing Automatic call summary and CRM field extraction are Phase 1 non-negotiables; extract facts from call transcripts is Phase 2 and rated High risk. They read the same source and do substantially the same thing. Either the Phase 1 items carry the Phase 2 legal gate, or the phasing needs correcting. This one is worth getting right — it is the only High-risk row in the document.
Rep coaching phasing The v3 amendment says "PULLED FORWARD"; the row still reads Phase 3. Low severity, but it is a governance-sensitive capability and should not be ambiguous.
Content that does not exist yet Objection library from real sales history, case studies tagged by vertical, approved competitive positioning, approved proof points, sales playbook, contract and proposal templates. All are Phase 1 dependencies. None has a named owner. Most likely to move dates This is DMA subject-matter time, not engineering time, and in my experience it is the most underestimated item in projects of this shape.
Pricing policy and guardrails The Canonical's own dependency list requires documented pricing policy before guardrails, negotiation limits or autonomous sending can be built. The guardrails are load-bearing throughout and never specified. The transaction cluster cannot be estimated or scheduled until this exists. It is DMA's commercial judgement; I can run the sessions that get it written down.
Compliance positions Consent basis and jurisdictions for contacting 30,000 dormant records; comfort with an agent binding the company and taking payment; scope of call-recording use; a data processing agreement covering records and transcripts passing through a model provider. A restrictive answer on outreach consent shrinks the contactable universe, which is the core value thesis. Needs an owner and a date, not a resolution from me — I do not give legal advice.
Acceptance criteria and baselines The Phase 1 scorecard names seven metrics and sets no targets. There is no baseline close rate, average deal size, cycle length or inbound volume to measure lift against. Determines how much tuning sits inside a fixed price. A revenue-based milestone gate carries materially more delivery risk than a functional one and should be priced as such.
13 / What we need from DMA

Dependencies, and when each one bites

These are the items I am assuming DMA provides or resolves, and which milestone moves if one of them slips. Short list, honest answer.

NeededByIf it slips
Nutshell API credential, ideally read-only firstBefore day oneEverything. This gates the whole project and two days of testing settles the largest unknown in it.
Agreement on what a good lead looks likeDay 10–15Day 20 scoring. The first pass will disagree with you; those sessions are where it gets fixed.
The approved list of permissions by field — what this system may write, and what stays people-onlyDay 10Nothing, if we have it. It is an hour of someone's time and it prevents the class of problem nobody forgives.
Escalation routing — which people or channel each of the seven triggers goes toDay 25Day 30, and only briefly. It is a short list you can change yourself afterwards; the one route that cannot be empty is "prospect asked for a human".
A second named approver for rule and guardrail changesDay 30Nothing directly, but without it the four-eyes control in section 09 is decorative.
Consent position on contacting dormant recordsDay 20Day 30 outright. Nothing goes out without it.
Sending domain and ESP accountDay 20Day 30. Warm-up takes calendar time that cannot be compressed by working harder.
Written pricing policy, discount floors, payment terms, deal ceilingDay 30The entire transaction cluster. This is the one I would start on today regardless of who you hire.
Approved contract and proposal templatesDay 35Day 45 transaction path. No capability in the Canonical produces these.
Objection library, proof points, case studies, competitive positioningRolling, from day 15Quality of every customer-facing message. Does not block a demo; absolutely determines whether the system converts.
Website platform access and chat widget decisionDay 30Day 45 inbound engine only. The reactivation engine is unaffected.
Zoom transcript access and retention scopeDay 30Zero-Entry CRM items at day 45.
E-sign and payment accounts in DMA's nameDay 35Day 45 staging walkthrough.
A named person for weekly reviewFrom day oneNothing on paper, everything in practice. This is worth considerably more than it sounds.

Why I sent you this instead of a demo

A generated proof of concept answers the question "can something that looks like this exist?" The answer is yes, and it has been yes for about two years. It does not answer the questions that decide whether this project works: whether your Nutshell account supports the writes the design assumes, where a rep looks when the AI has done something, what stops two agents writing to the same opportunity, how a discount floor is enforced rather than suggested, what a clever visitor can talk the website agent into doing, how you take back something it got wrong, and what happens to a signed contract that was never paid for.

These three documents answer those. They are not complete, and I have said clearly where they are not. Treat this as an early plan rather than a settled one: I have reviewed and shaped every position in it and will defend each one, and I would still expect Discovery to revise parts of it thoroughly. If you hire me, this is the shape of what I will build. The real design will be better than this one, because it will have your live Nutshell account, your pricing policy and your sales history behind it rather than my inferences about them.

Happy to walk any section of this on a call, and equally happy to be told I have read something wrong in the Canonical — it is a dense document and I would rather be corrected now.

Ryan Abrams · Robot Sea Monster Games
Prepared against the DMA AI Revenue System v3 Canonical and the finalist confirmation questionnaire.
Part Three sets out how I would build it.
Illustrative screens in this part; no live data anywhere.