Skip to main content

Release gates for maintainers

This page holds the release machinery from the public agent guide at flymy.ai/skill.md. A personal user connecting one account does not need to learn it, and an assistant building a requested worker never runs it: it creates and publishes only the agent the user asked for. The same lifecycle gate is described for humans in Create, Freeze & Embed.

Publication contract​

Publication contract: AGENTS.md is the canonical source. https://flymy.ai/skill.md and https://flymy.ai/AGENTS.md must serve its exact bytes. Base lifecycle live-verified on 2026-08-27. MCP resource-scope additions are a 2026-08-28 release candidate; run both early gates before promotion and repeat the clean-agent gate after production promotion. These candidate guide bytes are aligned to Agents MCP 0.7.0-rc.19 and Python SDK 1.2.0rc5. Capability-detect the installed package and live MCP schemas instead of assuming an older or newer artifact has this contract.

A release-candidate run must use an explicit candidate release contract, MCP, skill, and Agents API roots and must never fall back to production.

Release contract check​

Release maintainers must first fetch the exact guide origin's /agent-guide-release.json. Its compatible outer schema-v1 guide projection contains one release_contract with schema version 2. Before any product mutation, the release harness must require the intended candidate channel and coordinates, recompute the product bundle and candidate tuple hashes, compare the tuple to an independently approved release pin, and verify the pinned bytes for /skill.md, /llms.txt, the OpenAPI URL, /agents-mcp-tool-schema.json, and /embedded-capabilities.json. It must also verify the declared Agents MCP package and image, Python SDK wheel and minimum version, consumer identities, image digests, and lifecycle evidence state. Missing members, stale hashes, incompatible versions, blocked coordinates, or a candidate/stable mismatch stop the release gate. This tuple check is release machinery; a personal user connecting one account does not need to learn it.

Embedded release - two controlled customers​

For an embedded release, have the trusted release harness call the same deployment for two controlled customer identities after the assistant has shipped it. Normal product IDs come from authenticated server sessions; the harness may mint synthetic IDs only inside its trusted test identity store. Do not put those IDs in the model prompt or MCP tool arguments. Require distinct principal and execution IDs while keeping the same agent, frozen version, deployment, and billed owner. Replay each identical request with its original idempotency key and require the original execution ID.

Release-maintainer-only connectionless fixture​

This deterministic platform fixture is not part of an ordinary user's agent lifecycle. A release maintainer may run its one separately labeled disposable fixture agent once against an assembled candidate. An assistant already building a requested worker must skip this entire section and publish that same requested agent and compilation - never create a smoke agent beside it. The separate fresh-assistant gate also creates exactly one requested evaluation agent and must not add this fixture.

Direct-call support and embedded support are different. A configured connector can still fail embedded preflight. For a non-interactive acceptance test, the live-verified connectionless catalog adapter is hackernews, action HACKERNEWS_GET_LATEST_POSTS, with exact arguments:{}. Its schema lookup currently returns the discovered runtime under not_found; a direct call returns result.response_data.hits. Search with the exact query hackernews; broader prose currently ranks poorly. Resolve or add its integer account-tool ID, because available_tools accepts IDs, not slugs or runtime names:

: "${FLYMYAI_AGENTS_API_ROOT:?Set the exact Agents API root ending in /api/v1/agents}"
A="${FLYMYAI_AGENTS_API_ROOT%/}"
auth=(-H "X-API-KEY: $FLYMYAI_API_KEY")
: "${OPERATION_LABEL:?Persist a unique audit label before any write}"
: "${SOURCE_RUN_KEY:?Reserve a caller-owned key for the source run}"
tools=$(curl -fsS --max-time 30 "${auth[@]}" "$A/tools/?view=slim")
TOOL_ID=$(jq -er '.[]|select(.mcp_tool=="hackernews")|.id' <<<"$tools" || true)
if test -z "$TOOL_ID"; then
added=$(curl -fsS --max-time 30 -X POST "${auth[@]}" \
-H 'Content-Type: application/json' --data '{"mcp_tool":"hackernews"}' "$A/tools/")
TOOL_ID=$(jq -er '.id' <<<"$added")
fi

Create and run a real tool-attached owner agent through REST. The prompt names the exact discovered action so the frozen version retains the connector:

agent=$(curl -fsS --max-time 30 -X POST "${auth[@]}" \
-H 'Content-Type: application/json' \
--data "$(jq -nc --argjson tool "$TOOL_ID" --arg label "$OPERATION_LABEL" \
'{name:("Embedded Hacker News smoke ["+$label+"]"),user_prompt:"Call HACKERNEWS_GET_LATEST_POSTS once and return only the first post title.",available_tools:[$tool],output_schema:{type:"object",properties:{title:{type:"string"}},required:["title"],additionalProperties:false}}')" \
"$A/tasks/")
AGENT_ID=$(jq -er '.uuid' <<<"$agent")
run=$(curl -fsS --max-time 30 -X POST "${auth[@]}" \
-H "Idempotency-Key: $SOURCE_RUN_KEY" \
-H 'Content-Type: application/json' --data '{"variables":{}}' \
"$A/tasks/$AGENT_ID/run-loop/")
EXECUTION_ID=$(jq -er '.id' <<<"$run")
since=
for attempt in $(seq 1 150); do
if test -n "$since"; then
status=$(curl -fsS --max-time 30 -G "${auth[@]}" \
--data-urlencode "since=$since" "$A/executions/$EXECUTION_ID/status/")
else
status=$(curl -fsS --max-time 30 "${auth[@]}" \
"$A/executions/$EXECUTION_ID/status/")
fi
since=$(jq -r '.last_step_id // empty' <<<"$status")
if jq -e '.is_settled' <<<"$status" >/dev/null; then break; fi
sleep 2
done
jq -e '.status=="completed" and .error==null and (.result.title|type=="string")' \
<<<"$status" >/dev/null
freeze=$(curl -fsS --max-time 30 -X POST "${auth[@]}" \
"$A/compilations/freeze-instruction/$EXECUTION_ID/")
COMPILATION_ID=$(jq -er '.id' <<<"$freeze")

The final POST freezes this fixture's accepted EXECUTION_ID exactly once. Persist the returned COMPILATION_ID immediately and continue with the universal publish sequence in the Embedded and resale builder reference. If its response is lost, reconcile by the known execution ID rather than repeating it. Before publish, this fixture's access contract must contain a hackernews requirement with connection_required:false. After publish, the deployment run uses a synthetic external_user_id without a connection link, FlyMyAI account, or end-user key. Repeating the identical body and Idempotency-Key within the 2-day retention window must return the same execution ID.

Early release gate - one deployment, two customers​

Run two independent blocking gates against the release candidate before broad regression suites and before production promotion:

  1. A deterministic API gate validates the exact public release tuple and guide before mutation, then creates one labeled agent, performs one accepted live run, one real instruction freeze, one immutable version, and one deployment. It calls the deployment for two synthetic external_user_id values and proves idempotent replay.
  2. A separate fresh-assistant evaluation starts in an empty directory with only those exact guide bytes as FlyMyAI documentation. Give it a normal embedded-product request without lifecycle, customer-ID, or split hints. Grade structured MCP tool calls. After it publishes one deployment, the trusted harness exercises two customer bindings and independently verifies every ID over REST; prose claiming that calls happened is not evidence.

These are release-maintainer gates, not extra agents in a user's lifecycle. The deterministic gate owns its one disposable fixture agent. The fresh assistant owns exactly one requested evaluation agent. It must not also run the fixture section above. Before either gate, require all candidate coordinates:

: "${CANDIDATE_SKILL_URL:?Candidate guide URL is required}"
: "${CANDIDATE_RELEASE_CONTRACT_URL:?Candidate release contract URL is required}"
: "${CANDIDATE_RELEASE_TUPLE_SHA256:?Approved candidate tuple hash is required}"
: "${CANDIDATE_MCP_URL:?Candidate MCP URL is required}"
: "${CANDIDATE_AGENTS_API_ROOT:?Candidate Agents API root is required}"
: "${CANDIDATE_OPENAPI_URL:?Candidate OpenAPI URL is required}"
case "$CANDIDATE_AGENTS_API_ROOT" in */api/v1/agents) ;; *) exit 1 ;; esac
export FLYMYAI_AGENTS_API_ROOT="$CANDIDATE_AGENTS_API_ROOT"

The product tuple excludes the two benchmark result attestations so they can move from pending to passed without a self-referential hash. After both gates pass, pin the exact ready manifest bytes for the final read-only promotion check. A changed product tuple always requires both lifecycle gates again.

Do not substitute the production roots when any candidate coordinate is missing. The trusted release harness must mint and persist the two synthetic customer IDs in its own authenticated test identity store before any request. The assistant receives the product task, not authority to choose those IDs. The owner Agents MCP remains an authoring surface and does not gain arbitrary customer administration or a model-supplied external_user_id. Deployment publish, principal setup, customer connection, and customer execution remain typed REST and SDK operations performed by the trusted harness. A separately provisioned customer-bound MCP exposes only its fixed deployment/customer binding.

For an MCP resource-scope release, add a blocking scope contract before the general A/B topology checks. It must create several same-toolkit connection fixtures with distinct aliases and public IDs, place them in one revisioned owner set, grant that set directly and through one flat agent group, require an exact selector under ambiguity, and prove a revoke fails before provider dispatch. It must also create one external-principal named mapping and reject a foreign principal, stale revision, and a body that supplies both resource_set_id and connections. Run the real provider OAuth acceptance test separately after interactive authorization.

For both harness-owned IDs, require exactly one result from GET /external-principals/?deployment=<deployment_id>&external_user_id=<id> and different principal public_id values. For each customer execution, require exactly one GET /runtime-snapshots/?execution=<execution_id> result. Both snapshots must contain the same deployment and immutable agent_version, and each snapshot's billing_user must equal GET /me/ field user_id, while execution and principal differ. Replay each identical deployment request with its original Idempotency-Key and require the original execution ID. Read both prices with the same builder key. Query deployment access separately for both external IDs and reject any connection or binding UUID that crosses principals.

Use a connectionless supported adapter for this deterministic topology gate, then test each real provider separately. Preserve the exact prompt, guide bytes and SHA-256, structured transcript, operation ledger, created IDs, REST evidence, and cleanup result. Cleanup only labeled test resources: revoke any test connections, disable both principals, archive the deployment, then soft-archive the agent. Immutable versions, snapshots, prices, and other audit rows remain. Connection revocation can intentionally leave a revoked connection audit row and an empty binding row; require terminal revoked/unbound state, not physical deletion. A missing principal, one shared principal, a changed deployment/version, an idempotency replay that creates a new execution, owner-billing drift, incomplete cleanup, or a task status patch standing in for instruction freeze blocks release.

Resource-scope rollout​

Release maintainers deploy expansion and compatible readers before enabling multi-alias writers or new clients. Before removing compatibility readers or reversing the alias constraint, run the backend's read-only alias downgrade preflight and require a clean result. If it finds non-default aliases or scoped state, keep the compatible backend live and do not rewrite customer mappings to force a rollback.

Production resource observations​

These numbers explain the limits in the guide's resource-discipline rules:

  • Poll /executions/{id}/status/?since=. A live settled response was about 1.1 KiB; full detail for a tiny exploratory run reached about 67 KiB per poll.
  • Avoid account-wide agent, compilation, full catalog, and price lists in request paths. Several are unpaginated and grow with the account.
  • Price detail is O(tool calls + LLM turns). A tiny live response was 759 bytes, but long agents grow linearly and can require rate lookups.
  • Embedded access was about 6.7-9.7 KiB and reads deployment, version, requirements, principal, connections, and bindings. Poll only during setup with backoff.
  • Bound concurrency and timeouts. Sync or streaming model inference can retain an async DB session plus request and upstream network streams. An arbitrary custom MCP call occupies a synchronous worker thread while waiting on the remote server.
  • Never use unbounded asyncio.gather. Use a semaphore, queue, and per-user and global quotas.
  • Cap files before decode and stream large payloads. Reconcile ambiguous external writes before retrying.

Production has seen 354,000-token append-message responses and 195,000-token agent lists. One worker serves 30 threads and has been killed around 1.4-2.0 GiB under a 2300 MiB limit. One oversized response can drop all 30 requests. After deployment, inspect pod working-set memory, OOM kills, restarts, latency, DB query count, and response sizes.