Stimulir Quick Start

From your first trace to a measured improvement

Install Stimulir, connect your application or import its existing traces, then use a coding agent to inspect, evaluate, and improve one AI behavior safely.

The workflow

  1. Connect a project and route inference through Stimulir, or import existing trace evidence.
  2. Observe the prompt, model, tags, cost, and latency behind each response.
  3. Preview a privacy-checked, immutable evaluation cohort before spending.
  4. Compare the incumbent with one controlled prompt or model change.
  5. Review comparable evidence. Promotion is always a separate human action.

1. Install the CLI and coding-agent skills

bash
# Recommended isolated CLI installation
uv tool install stimulir

# Or install with pip
pip install stimulir

# Add the lifecycle skills to Codex or Claude Code
npx skills add stimulir/skills

stimulir --version

The CLI provides authenticated platform commands. Skills give a coding agent the safe workflow, so you ask for an outcome rather than manually reproducing capture, privacy, evaluation, and monitoring steps.

SkillUse it for
connectVerify project credentials and one cost-visible inference call.
migrate-inferenceMove direct provider calls onto Stimulir.
prompt-versioningCreate durable prompt versions and labels.
capture-tracesCurate observed traffic into data assets.
eval-runMeasure one controlled prompt, model, or adapter change.
rsiRun Improve: inspect, preview, run, and review a durable diagnostic hill climb (legacy skill package name).
eval-promoteApply one reviewed proposal with explicit authorization.
usage-auditReview task cost and billing evidence.

2. Connect an application

In the Stimulir Console, create or select a project and create a productionhyb_* key. Store these values in the application repository, not in source control.

.env
STIMULIR_API_BASE=https://api.stimulir.com
STIMULIR_API_KEY=hyb_...
STIMULIR_PROJECT_ID=...
bash
uv add stimulir   # or: pip install stimulir
python
from stimulir import StimulirClient

client = StimulirClient()
response = client.agent(
    model="<your-enabled-provider/model>",
    input="Summarise the customer request.",
    metadata={
        "prompt_key": "support.summary.agent",
        "prompt_version": 1,
        "prompt_label": "production",
    },
    tags=["support", "summary"],
)
print(response.output_text)

Replace the model with a route enabled for your project and use an input from your own workflow. The support prompt key, version, label, and tags in this example are illustrative: use your actual registered prompt lineage and workflow tags so the trace can be matched to the behaviour you want to evaluate.

You are connected when the call succeeds and its trace appears under Engineering → Traceswith the correct project, prompt lineage, model, tags, latency, tokens, and cost where available.

An application key is already workspace-bound. Supply the project ID; do not add a workspace ID.

3. Choose your evidence

SourceUse it whenWhat it proves
Production tracesYour application already routes inference through Stimulir.How the deployed prompt and model behaved on real traffic.
Trace importYou have Langfuse, LangSmith, Arize Phoenix, JSON, JSONL, or CSV evidence.How a chosen prompt/model performs on imported cases.

Production traces

Capture the exact prompt_key, prompt_version, provider/model, workflow tags, input, output, and operational measurements. Improve will not map an unrelated tag to a prompt or invent missing lineage.

bash
stimulir lab improve inspect --env-file .env

# Narrow a real workflow without starting a run
stimulir lab improve inspect \
  --trace-tag <your-workflow-tag> \
  --source-window 30d \
  --count 50 \
  --env-file .env

Import existing traces

Open Engineering → Traces → Import traces and upload JSON, JSONL, or CSV. Stimulir recognises common Langfuse, LangSmith, and Arize Phoenix export fields, preserves source IDs, tags, and prompt lineage when present, and skips duplicate rows when the same export is retried. Review the imported rows, then send the useful selection to Data for curation and snapshotting.

Import is currently a Console action. A coding agent can inspect a local export, identify missing lineage, and guide you through the import; it must not claim it uploaded the file unless a supported import command exists.

4. Work through Codex or Claude Code

Open the coding agent in the application repository so it can resolve the project's .env.

Use the other lifecycle skills

OutcomeAsk the coding agent
ConnectConnect this application to Stimulir and verify one real inference call.
InspectShow available production workflows, tags, prompts, models, date ranges, and counts. Do not start anything.
Prepare an importInspect the Langfuse JSONL export in ./data, report missing prompt lineage, and tell me how to import it through Engineering → Traces.
EvaluateEvaluate prompt v4 against the newest 50 eligible support traces. Do not promote.
Review spendAudit cost per task for this project over the last 30 days.

5. Run Improve

See Improve for the complete CLI, SDK and controller reference.

Improve is a durable diagnostic workflow, not a generic evaluation prompt. Give the coding agent the cohort, experiment boundary, managed-inference permission, monitoring requirement, and the evidence you expect back. The skill will inspect first, produce a signed read-only preview, start once when runnable, and monitor the server-owned run to a terminal state.

Copy this prompt for production traces

text
Run Improve on the newest 100 eligible traces for <workflow or exact tag>, or use the full eligible cohort when fewer than 100 are available.

Use the trace-derived production prompt and model as the incumbent. Use prompt-only mode and one iteration. Use evaluator, judge, and proposer routes enabled for this project. Show the recommended configuration and estimated spend before running; ask before enabling managed inference or changing the approved budget.

Inspect first and show me the resolved prompt, model, trace count, evaluator, judge, proposer, and nearest runnable alternative for any blocker. Keep the trace-derived model as the incumbent. Resolve a typed blocker automatically when there is one safe, exact, authorized alternative; ask me only when authorization is missing or materially different choices remain. Preview the exact cohort, then start automatically if it is runnable without asking me to reconfirm.

If the recorded prompt or inference lineage is unavailable, report the instrumentation or import-mapping gap and stop before spending budget.

Monitor the durable Improve run until it is completed, needs_input, or failed. Do not invoke another evaluation, capture, privacy, or Lab skill. Never promote automatically.

At completion, report the diagnosis; incumbent-versus-candidate quality, pass rate, latency, and cost; total measured spend and accounting confidence; rejected candidates; stopping reason; your recommendation to retain, reject, iterate again, or prepare a human-reviewed promotion proposal; the Improve run ID; and the Console link. Only prepare a proposal when the evidence supports it, and leave the production label unchanged until a human separately confirms promotion.

Use traces from your own application, captured through Stimulir or imported from your existing tooling. Confirm the recorded prompt and model before running. If lineage is missing, repair instrumentation or the import mapping first. Choose a representative cohort for your task; a small run can check setup, but promotion requires sufficient comparable evidence.

If the requested setup is blocked, the coding agent should lead with one recommended runnable alternative. It should explain the evaluator, judge model, and proposer as separate roles instead of asking for raw route details without a recommendation.

Preview the exact cohort

Replace every angle-bracket placeholder with values returned by inspection for your project. Use an evaluator suited to your task and an available judge route; the same selectors must be used for the run.

bash
stimulir lab improve preview \
  --source-window 7d \
  --trace-tag <your-workflow-tag> \
  --count 100 \
  --prompt-key <your-prompt-key> \
  --evaluator-key <your-evaluator-key> \
  --evaluator-version <your-evaluator-version> \
  --judge-provider <your-judge-provider> \
  --judge-model <your-judge-model> \
  --env-file .env

Review selected and eligible counts, privacy exclusions, incumbent prompt/model, evaluator and judge, experiment plan, duration estimate, accounting status, and typed blockers. A runnable preview returns a signed token and cohort fingerprint.

Start once, then let the server own the loop

bash
stimulir lab improve run \
  <the same reviewed selector> \
  --cohort-fingerprint <fingerprint> \
  --preview-token <token> \
  --idempotency-key <stable-retry-key> \
  --env-file .env

Never print the preview token. If the selector, evaluator, model, or experiment changes, preview again. Promotion remains disabled.

The server owns cohort creation, evaluation, diagnosis, candidates, accounting, stopping, and recovery. If asked to watch, the coding agent may use a read-only watcher and contextualise the terminal result; it must not silently steer, resume, or promote.

bash
stimulir lab improve status <run-id> --env-file .env
stimulir lab improve overview <run-id> --env-file .env
stimulir lab improve results <run-id> --env-file .env

Review quality, cost, latency, coverage, diagnosis, rejected candidates, stopping reason, and accounting completeness. Missing usage is unpriced, not zero. An incomplete or incomparable run is not a win.

6. Operate and get help

  • Engineering: inference, keys, BYOK, prompts, traces, privacy, data, and usage.
  • Lab: evaluations, Improve, adapters, RL, Doc-to-LoRA, and projectors.
  • Compute: GPUs, workers, storage, sandboxes, tunnels, and runtime secrets.
bash
# Authoritative flags for the installed version
stimulir --help
stimulir lab improve --help

# Cost evidence
stimulir usage summary --env-file .env
stimulir usage billing --env-file .env

Still blocked? Send the project ID, command, typed error code, and run ID. Never send an API key or preview token. Contact hello@stimulir.com.