Skip to content

Checkout sandbox runtime: remaining requirements (app-runtime image, COI / code-on-incus, isolation, lifecycle) #489

Description

@TonsOfFun

Follow-up to #477 / #479 (GitHub connections + app_runtime checkout sandboxes) and #478 / #482 (Claude Code connection). The engine and the platform's Incus backend (activeagents/activeagents, "Wire GitHub connections and checkout sandboxes into the platform") now agree on a contract. This issue tracks what still has to exist before a checkout sandbox runs real code in production.

What already works

  • The engine validates the repository against the owner's GitHub selection and gives the backend sandbox_session.checkout_spec (repository, ref, clone_url, username, token) and sandbox_session.runtime_environment (CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY).
  • IncusSandboxService (platform) creates the container from sandbox-app-runtime. It then:
    • fetches the ref into /workspace/app; the token rides only in the exec environment and an ephemeral http.extraHeader, never in instance config, argv or .git/config;
    • runs sandbox-app-boot /workspace/app with the Claude Code environment;
    • waits for :8080/up;
    • reads /workspace/runtime.json → { "mcp_path": …, "mcp_token": … }.
  • The session then serves as the MCP server sandbox:<session_id> for the owner's agents and evaluations.

Remaining requirements

1. The sandbox-app-runtime image and boot contract

  • Build and publish the Incus image sandbox-app-runtime (and -gpu). It needs git, Ruby via a version manager that honours the checkout's .ruby-version / .tool-versions, Node/Bun for JS builds, a local Postgres/SQLite, libvips, and Chromium if Playwright tools are expected.
  • Ship /usr/local/bin/sandbox-app-boot APP_DIR:
    • bundle install / JS install with caches;
    • bin/rails db:prepare against a throwaway database;
    • start the server on 0.0.0.0:8080 detached;
    • create a dashboard API key in the booted app for its MCP facade;
    • write /workspace/runtime.json (mcp_path = the checkout's engine mount + /mcp, mcp_token = that key);
    • exit non-zero with a useful stderr on failure. A convention such as bin/sandbox-setup or .activeagents/sandbox.yml in the repo would let an app customize the steps.
  • Decide which environment the booted app gets. For example: RAILS_ENV=development or a dedicated sandbox env, SECRET_KEY_BASE generated per sandbox, and provider keys only when the owner opts in.
  • Document the contract in docs/framework/dashboard.md next to the existing backend section.

2. Claude Code sessions in the sandbox: COI (code-on-incus) or similar

The Claude Code credential reaches the booted app, but nothing runs a Claude Code session against the checkout yet (the dashboard assistant still says "COI execution … not implemented"). Evaluate COI (code-on-incus), which runs Claude Code in isolated Incus containers, against a small runner of our own:

  • Choose how sessions run: COI driving Claude Code inside the checkout container, a sibling container sharing /workspace/app, or claude -p headless invoked through Incus exec.
  • Define the dashboard surface: start a session with a prompt, stream its transcript (Action Cable, like sandbox runs), and show the resulting diff.
  • Record sessions as traces/runs so they show in Interactions and can be evaluated.
  • Credentials: CLAUDE_CODE_OAUTH_TOKEN from runtime_environment only, never persisted in instance config, and scrubbed from transcripts.
  • Once this works, flip the assistant's connections.coi / connections.claude_code to supported.

3. Isolation and security

  • Network egress policy for app_runtime: allow GitHub, package registries and the chosen LLM provider, deny the platform's internal network and cloud metadata endpoints.
  • The container must not reach the dashboard's database or internal services. The runtime MCP endpoint should be reachable from the dashboard only (Incus network ACL / bridge rules).
  • Replace the user OAuth token with GitHub App installation tokens, per repository and short-lived. That limits a leaked token to one repo for about an hour.
  • Resource limits for app_runtime (currently the cpu_medium tier by default): disk quota for bundle install, max boot time (BOOT_TIMEOUT = 900s), process limits.
  • Review the AppArmor profile in sandbox-restricted for a full app (it was written for the Playwright and terminal images).

4. Lifecycle and UX

  • Async provisioning feedback: booting a checkout takes minutes, so the Settings "Start sandbox" button should follow the sandbox channel instead of the create response.
  • Surface live sandbox:* runtimes in the agent editor's MCP server picker and the Tools roster (Dashboard: connect GitHub via OAuth, choose repositories, and run agents/evals in a sandbox built from the checkout #477 follow-up), so the key doesn't have to be copied.
  • Evaluation runner option: "run this evaluation against sandbox X" without editing the agent's mcp_servers.
  • Re-sync a sandbox to a new ref (fetch + restart) instead of a new container. Expire and clean up containers and their runtime MCP tokens.
  • Publishing: push a branch or open a PR from a sandbox's working tree after a Claude Code session. This needs contents:write / PR scopes, which is an argument for GitHub App tokens.

5. Other backends

  • KubernetesSandboxService and CloudRunService (platform) do not handle checkout_spec yet. Port the same contract (init container for the fetch, runtime.json from a shared volume), or declare app_runtime Incus-only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions