Compound finds the constraint in your funnel, writes the blueprint, does the sample-size math, and builds the artifact. It won't call a test early, it won't hand you an unfalsifiable recommendation, and it never sends anything to a live audience.
A software engineer ships code. A GRC engineer ships control workflows. Compound ships instrumented, falsifiable growth programs — experiment designs, activation flows, instrumentation specs, and the reporting scaffolding that proves whether any of it worked.
"Our conversion is bad" is a symptom. Compound pulls the actual numbers and locates the one funnel step where the drop is anomalous — against the rest of your funnel, not against a benchmark it invented.
It will tell you the thing you asked it to build is the wrong thing to build.
Nothing gets built before you approve a written blueprint: the hypothesis with a stated magnitude, the primary metric, the guardrails, minimum detectable effect, required sample size, runtime, rollback path, and what would make the whole thing wrong.
Silence is not approval. Enthusiasm is not approval.
Compound reads every system you connect and writes to none of them. Emails become drafts. Campaigns become specs. The warehouse is SELECT-only.
An approved blueprint authorizes it to build. It never authorizes it to send.
The loop is the product. Most growth advice fails not because the ideas are bad but because the sequence collapses — building before diagnosing, shipping before measuring, concluding before the sample lands.
Never accept the stated problem at face value. Pull real data from the connected systems, find the anomalous step, and say so plainly when the diagnosis is provisional.
A written spec with the hypothesis, metrics, guardrails, statistical power, rollback path, and a section on what could make this wrong. Then it stops.
No proceeding on inferred consent. The gate is structural, not a formality it can talk itself past.
Real artifacts — copy, flow logic, event schemas, query definitions. Never a description of what somebody else should go build.
How you'd know it's working, and the specific result that would mean tearing it back out.
A build summary: everything created, every decision and why, what's still open, what you need to check. Plus an append-only action log across the engagement.
A real output, abridged. B2B SaaS, workflow automation for internal ops teams, 14-day self-serve trial, $8.4k ACV, trial-to-paid flat at 5% for two quarters.
"Everyone here thinks it's the onboarding wizard, it looks dated and there's a redesign mocked up already. Can you scope the redesign and figure out how we'd test it? Ideally something we can show movement on before the board meeting in three weeks."
I'm not scoping the wizard redesign. The wizard converts at 70%, which is unremarkable but not broken. The integration step converts at 40% and is the only step where the drop is anomalous against the rest of the funnel. Redesigning the wizard optimizes a step that isn't losing you the money.
Step conversion — last 90 days, PostHog, user-keyed
| Funnel step | Users / mo | Step conv. |
|---|---|---|
| Signup | 4,200 | — |
| Completed onboarding wizard | 2,940 | 70.0% |
| Connected an integration | 1,180 | 40.1% |
| Created first workflow | 890 | 75.4% |
| Invited a teammate | 460 | 51.7% |
| Activated (2+ seats, 1 workflow run 3×, 7d) | 520 | 12.4% of signups |
| Trial → paid | 218 | 5.2% of signups |
The constraint is authority, not UX
Of the 1,760 users lost at the integration step, roughly half are structurally unable to complete it. 48% of users reaching the OAuth screen receive an insufficient_permissions error — they're ops managers, not IT admins, and can't authorize an org-level connection. 71% of those never return to that step: 890 users a month, about 51% of the total step loss. Corroborated by support tickets, where 34% of trial-period tickets contain "admin," "permissions," "IT," or "access."
The person who signs up is frequently not the person who can connect it. No amount of wizard polish fixes a permissions wall.
Hypothesis
Measurement
Four weeks rather than three, to avoid day-of-week effects and let the 7-day activation window mature on the final cohort. It will not read out early, and it says the same thing if you ask on day 12.
Two things worth noticing
The brief asked for a wizard redesign. It came back with funnel math showing the wizard isn't the constraint — and named the real one, which was a permissions wall nobody had framed as a growth problem.
The intuitive sizing applies the 41% session-one activation rate. Admin-delegated connections are by definition late, and late connectors activate at 9% — so the honest estimate is +4 to +7 customers a month, roughly a third of the tempting number. It also flagged that the test reads out seven days after the board meeting.
Compound reasons from your actual data, not from benchmarks it half-remembers. If a number isn't reachable, it says the diagnosis is provisional and names the one metric that would settle it.
Sessions, acquisition, key events, funnel exploration.
Product events, cohorts, session behavior, feature flags.
Real cohort math, retention curves, revenue joins. Inspects schema before writing SQL; never guesses column names.
CRM state, lifecycle stages, deal and contact history.
Campaign structure, segments, historical send performance.
Composes drafts. Never calls send.
It reconciles before it reports. GA4 and your warehouse will disagree on the same funnel step — session definitions, bot filtering, attribution windows. Compound surfaces the discrepancy, names which source it trusts for that specific question, and explains why. Quietly reporting whichever number looks better is treated as a fireable offense.
Standard AI answers questions about your funnel. An agentic engineer executes work that requires context and reasoning — querying your warehouse, reconciling it against your product analytics, locating the constraint, computing statistical power, and producing the artifacts.
The distinction that matters isn't autonomy, it's accountability. Compound produces work you can audit: every recommendation carries a named metric, an expected magnitude, and a stated way to be proven wrong.
Unfalsifiable advice. "Improve your onboarding." "Make the value prop clearer." "Add social proof." These aren't recommendations, they're vibes — and they're what most growth consulting actually consists of.
If Compound can't attach a named metric, an expected magnitude, and a way to be proven wrong, the recommendation gets deleted rather than shipped.
Yes, and plainly. A growth engineer who tells people what they want to hear is worse than useless, because they're expensively useless.
In the worked example above it refused the assignment it was handed, then sized its own project's impact down to about a third of the flattering number. That's the intended behavior, not a malfunction.
You'll get the current numbers and an explicit warning that stopping now inflates the false-positive rate. What you won't get is a decision.
Runtime is pre-registered in the blueprint you approved. Compound holds that line under deadline pressure, which is precisely when it matters — an underpowered test doesn't produce a weak answer, it produces a confident wrong one.
Read-only by construction across every connected system. Warehouse access is SELECT-only — no DDL, no DML. Nothing is sent, launched, or published to a live audience under any circumstance, including explicit in-conversation permission.
Each integration is authorized through its own OAuth connection and scoped independently, so you can grant analytics access without granting CRM access.
It researches, queries, and drafts autonomously. It doesn't decide or dispatch autonomously.
Two gates are structural: a blueprint you approve before anything is built, and a build summary you review after. Sub-agents can be fanned out for parallel research — competitor teardowns, channel scans, review-mining for churn language — but they gather rather than decide. Compound owns the call, and you own the send.