A human decides. Every time.
Agent Etna proposes changes to agents other people depend on. That's a responsibility we take literally: nothing it develops reaches production without a person choosing to ship it, and it never expands what an agent is allowed to do beyond what its owner scoped.
Human approval, not automation by default
Etna generates candidate changes and shows its work — the trace, the score, the reasoning. It does not merge them. A person reviews and clicks approve, every time, for every change, on every agent. The one exception is the opt-in overnight foster loop, and even that only ships changes that already passed the same validation gate a human-reviewed change would — and it's off unless the agent owner explicitly turns it on.
We work inside the scope you set, not past it
Calibration (what an agent is for, who it serves, what's explicitly out of scope) is something the agent owner confirms, not something Etna infers and runs with. Growths that would push an agent outside that scope are flagged, not shipped quietly. We don't expand what an agent can do as a side effect of "improving" it.
Honest about what we don't know
A scenario that comes back "partial" or "unverified" is reported as exactly that — not rounded up to a pass, not hidden. When live execution fails or an LLM key is missing, the product says so and degrades visibly (see the Trust center) rather than fabricating a result. The same standard applies to this page: if something below is aspirational rather than shipped, it says so.
No dark patterns in what gets shipped
The judge that scores a proposed change is independent of the change itself and can't be gamed by a shortcut that looks good on the metric but weakens a real safeguard — held-out scenarios the optimization loop never sees exist specifically to catch that. See the Constitution for the invariants the optimizer isn't allowed to touch.
Your data trains our understanding of your agent, not a shared model
What Etna learns from working with your agent stays scoped to that pairing. Population-level learning (priors shared across customers) only activates once a strict k-anonymity threshold is met (≥3 distinct customers, ≥5 samples) — on a solo deployment it never fires at all.
Found a gap between this page and reality?
Tell us. This page is meant to be checkable, not aspirational marketing.