Skip to content

feat(cli): add identity-aware sandbox deletion waiting #3944

Description

@elezar

User Story

As an automation author, I want the OpenShell CLI to wait for the specific sandbox I deleted to disappear, so that I can continue safely without parsing human-readable errors or confusing a same-name replacement with the original sandbox.

Problem Statement

openshell sandbox delete can return after the gateway accepts asynchronous deletion while the sandbox record still exists. The deletion response contains the original sandbox_id, but the CLI currently discards that identity and offers no deletion-wait workflow.

Callers that need to confirm deletion must currently poll sandbox get or enumerate sandbox list. Polling get is unsafe for CLI automation because all command failures currently have the same generic nonzero process status, so NOT_FOUND cannot be distinguished reliably from authentication, transport, or server failures. Enumerating list is inefficient and tests name presence rather than the identity that was deleted.

PR #3792 demonstrates this gap in the sandbox lifecycle conformance scenario. Issue #2040 tracks machine-readable CLI errors, but structured errors alone would still leave each caller to reproduce identity-aware deletion semantics.

Impact / Why This Matters

Automation can report deletion as complete after an unrelated CLI failure, repeatedly enumerate every sandbox, or wait on a newly created sandbox that reused the deleted sandbox's name. These workarounds are fragile, duplicate SDK behavior, and increase gateway traffic during polling.

The Rust, Python, and TypeScript SDKs already support identity-aware deletion waiting. The CLI should expose equivalent behavior so scripts and conformance scenarios do not need to reimplement it.

Proposed Design

Add an opt-in deletion wait to the CLI:

openshell sandbox delete <name> --wait --timeout 60s

When --wait is set, the command uses the sandbox_id returned by DeleteSandbox as an opaque identity guard and does not return success until the original sandbox is gone. It succeeds when the named lookup returns NOT_FOUND or resolves to a different sandbox identity. Authentication, authorization, transport, and server failures do not establish deletion and must be surfaced to the caller. A timeout returns a nonzero exit status with a clear diagnostic.

Without --wait, preserve the current accepted-deletion behavior. The immutable ID remains an implementation detail; users continue to address sandboxes by canonical name.

This issue does not require a public get-by-ID command. It may be implemented by reusing the existing SDK deletion-wait semantics or an equivalent CLI-internal mechanism.

Acceptance Criteria

  • openshell sandbox delete <name> --wait waits until the specific sandbox identity returned by deletion no longer resolves.
  • The wait succeeds when the sandbox lookup returns NOT_FOUND.
  • The wait succeeds when the name resolves to a different sandbox ID, without deleting or waiting for the replacement.
  • Authentication, authorization, transport, malformed-response, and server failures never count as successful deletion.
  • --timeout bounds the wait and timeout failure is machine-detectable through a nonzero exit status.
  • Existing behavior remains available when --wait is omitted.
  • Completed, accepted, already-absent, same-name replacement, timeout, and non-NOT_FOUND failure paths are covered by tests.
  • CLI documentation explains the waiting and identity semantics.

Alternatives Considered

Treat any nonzero sandbox get result as absence. This produces false success for authentication, transport, and server failures.

Poll every page of sandbox list. A successful list distinguishes absence from command failure, but collection enumeration is inefficient and a name-only check can follow a replacement sandbox.

Add a public get-by-ID command. Direct ID lookup would distinguish identities, but OpenShell's public API convention uses canonical names and keeps immutable IDs at internal boundaries. Comparing the ID returned by a named lookup with the deletion result provides the needed guarantee without expanding the public reference model.

Expose only machine-readable CLI errors through #2040. That is independently valuable and enables safer single-resource polling, but callers would still have to implement timeout, retry, and same-name replacement behavior themselves.

Agent Investigation

  • DeleteSandbox already returns a typed deletion outcome and the original sandbox_id when it finds a target.
  • The Rust SDK's wait_deleted helper completes on NOT_FOUND or when the same name resolves to a different ID, and propagates other errors.
  • The CLI currently prints the deletion outcome but discards sandbox_id.
  • Kubernetes follows the same identity model: kubectl delete records the deleted resource UID and treats NotFound or a different UID as completion while waiting.

Related: #2040, #3051, PR #3792.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:cliCLI-related work

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions