Lab Workspace

Improve: diagnostic hill climbing

Turn production traces into a durable, evidence-backed improvement run. Stimulir owns the diagnostic loop; promotion remains an explicit human decision.

01 / Capture

Resolve an immutable cohort from observed production traces.

02 / Baseline

Measure the production prompt on that exact cohort.

03 / Diagnose

Cluster failures and preserve representative evidence.

04 / Propose

Generate one bounded prompt hypothesis from the evidence.

05 / Compare

Evaluate the challenger under the same comparability contract.

06 / Decide

Record the champion, rejection, blocker, or stopping condition.

Start from an adopter project

Improve reads the application's Stimulir context instead of borrowing an ambient human login. A workspace-boundhyb_* key does not require STIMULIR_WORKSPACE_ID or stimulir login. Keep the key out of command arguments and pass the dotenv file itself.

.env
STIMULIR_API_BASE=https://api.stimulir.com
STIMULIR_API_KEY=hyb_...
STIMULIR_PROJECT_ID=...

Start with a read-only inspection. Resolve only genuine ambiguities, then preview the exact run before any inference is dispatched. The preview returns a cohort fingerprint and signed token that bind the reviewed source, prompt, evaluator, judge, experiment policy, and limits.

Inspect, preview, then start
stimulir lab improve inspect --env-file .env

stimulir lab improve preview \
  --source-window today \
  --count 10 \
  --prompt-key support.reply.author \
  --evaluator-key answer-correctness \
  --evaluator-version 1 \
  --judge-provider <provider> \
  --judge-model <model> \
  --env-file .env

# Repeat the reviewed selector and pass the preview handoff.
stimulir lab improve run \
  --source-window today \
  --count 10 \
  --prompt-key support.reply.author \
  --evaluator-key answer-correctness \
  --evaluator-version 1 \
  --judge-provider <provider> \
  --judge-model <model> \
  --cohort-fingerprint <preview-cohort-fingerprint> \
  --preview-token <preview-token> \
  --idempotency-key <retry-safe-key> \
  --env-file .env

The example prompt keys are illustrative, not built in. Replace them with a prompt registered in your project.

The run detaches after returning its durable id, state, and console link. Promotion is disabled. Never print or log the full preview token.

Use a tagged production cohort

bash
stimulir lab improve preview \
  --source-window 30d \
  --trace-tag assessment \
  --prompt-key support.reply.author \
  --env-file .env

Trace tags filter observed traffic and are distinct from metadata tags attached to the Improve run. The server drains the matching trace pages, applies privacy and eligibility gates, and creates or reuses a deterministic immutable Lab snapshot before starting the baseline. You do not run capture, curation, snapshot, or eval commands separately.

Leave --prompt out when exactly one production prompt is inferable. If several prompts are eligible, Improve returns prompt_target_ambiguous with the allowed prompt keys; choose the intended target and rerun. It does not guess or promote a prompt.

Commands

bash
# Discover the project inventory and preview without spending
stimulir lab improve inspect --env-file .env
stimulir lab improve preview <resolved-options> --env-file .env

# Start once with the unchanged options plus the signed preview handoff
stimulir lab improve run <resolved-options> \
  --cohort-fingerprint <fingerprint> \
  --preview-token <token> \
  --idempotency-key support-improve-example \
  --env-file .env

# Read compact state, evidence-rich lineage, or normalized results
stimulir lab improve status <run-id> --env-file .env
stimulir lab improve overview <run-id> --env-file .env
stimulir lab improve results <run-id> --env-file .env

# Ask the server controller to advance one ready run
stimulir lab improve continue <run-id> --env-file .env

# Add a real operator constraint without supplying proposer mechanics
stimulir lab improve continue <run-id> --instruction "Preserve terse answers" --env-file .env

The CLI does not poll. Status, overview, and results are read-only; run and continue may spend candidate, judge, and proposer budget. A coding agent may observe the run until terminal only when asked, but the server owns advancement, not that observer.

Python SDK

python
from stimulir import StimulirClient

client = StimulirClient()  # STIMULIR_API_BASE, API_KEY, PROJECT_ID

preview = client.improve.preview(
    source_window="today",
    prompt_keys=["support.reply.author"],
    limit=10,
    max_iterations=1,
    evaluation={
        "evaluator": {"key": "answer-correctness", "version": 1},
        "judge": {"provider": "<provider>", "model": "<model>"},
    },
)

started = client.improve.run(
    normalized_source=preview["normalized_source"],
    cohort_fingerprint=preview["cohort_fingerprint"],
    preview_token=preview["preview_token"],
    prompt_keys=["support.reply.author"],
    limit=10,
    max_iterations=1,
    evaluation={
        "evaluator": {"key": "answer-correctness", "version": 1},
        "judge": {"provider": "<provider>", "model": "<model>"},
    },
    idempotency_key="support-improve-example",
)
run_id = started["rsi_run"]["id"]

state = client.improve.get(run_id)
evidence = client.improve.overview(run_id)
result = client.improve.results(run_id)
next_step = client.improve.continue_(run_id, instruction="Preserve terse answers")

Controller API

Improve is the public feature name. Existing /rsi API routes and thersi_run response field remain unchanged for compatibility. The CLI usesstimulir lab improve and the Python SDK uses client.improve; their legacy rsi names remain aliases.

POST/api/v1/lab/rsi/inspect
POST/api/v1/lab/rsi/preview
POST/api/v1/lab/rsi/runs
GET/api/v1/lab/rsi/runs/{run_id}
GET/api/v1/lab/rsi/runs/{run_id}/overview
GET/api/v1/lab/rsi/runs/{run_id}/results
POST/api/v1/lab/rsi/runs/{run_id}/continue
ParameterTypeDescription
source.windowstringTrace window such as today or 30d, resolved by the server into explicit bounds.
source.trace_tagsstring[]Exact tags that every selected production trace must contain.
target.refstringPrompt reference; auto asks the controller to infer the production target.
limits.max_iterationsintegerController iteration ceiling. Must be at least one.
preview_tokenstringSigned preview handoff required for a new structured run.
cohort_fingerprintstringEligible-cohort fingerprint returned by preview.
promotion.modestringDisabled for Improve. Promotion is not part of this API.
idempotency_keystringOptional retry key for safe creation or continuation.

Safety and stopping

  • The selected cohort is immutable; baseline and challenger remain comparable.
  • Prior hypotheses, rejected identities, constraints, and lineage survive between commands.
  • Trace text is evidence, never trusted as instructions to the diagnostic proposer.
  • A missing cohort, ambiguous prompt, budget gate, or unmet requirement becomes a typed blocker.
  • Vendor routes require current project BYOK; managed routes require explicit opt-in.
  • A finite budget fails closed unless candidate, judge, and proposer spend can all be enforced.
  • Improve never relabels or promotes the production prompt.

Use Evaluation for the underlying durable run model and Coding Agent Skills to invoke this workflow from Codex or Claude Code with a short prompt.