Create, run, and improve agents in one place. Sikaru manages the infrastructure and learning. You bring the judgment.
Fig. 01 · an example agent across 412 production runs. Reviewed changes are tested against the same cases; individual runs can still fail.
A frontier lab for your judgment.
You know the work. We build the machinery to make agents better at it—from the harness that runs them to the evals and training that improve them. Your expertise sets the direction.
Create an agent or bring your own. Connect its tools and knowledge. Sikaru handles the harness, deployment, memory, and tracing—then connects real work to evaluations and improvement.
Start from scratch, or move an agent you already operate.
Corrections, feedback, failures, recovery, traces. Sikaru learns from all of them.
Show what good looks like through feedback or examples. Sikaru proposes checks; you refine and approve them.
Execution, model calls, tools, and persistent memory, kept running and maintained for you.
Tests built from your real runs and feedback. Compare changes on the same tasks and catch regressions.
Reviewed changes to context and harness. Model training is in enterprise beta.
Feedback becomes signals. Signals become experiments. Sikaru tests what works; you guide the learning and approve what ships.
Sikaru runs your agent, connects its tools, and keeps the full trajectory. User feedback stays connected to the work that produced it.
Fig. 02 · from production feedback to a tested, approved change.
Sikaru improves what the agent knows, how it works, and the model behind it, where the work shows it should. Start with context and harness improvements; model training is in enterprise beta.
What the agent knows and is told.
How the agent runs the work.
The model's own behavior, shaped from your work.
Fig. 03 · context and harness improvements today. Model training and generated RL environments are in closed enterprise beta.
Improvement you cannot see is not improvement you can trust. Every change keeps its reason, the check it had to pass, and the name of the person who approved it.
Fig. 04 · an example learning release, from miss to shipped.
Your work stays in your workspace. We do not train foundation models on it.
Nothing reaches the agent until a check passes and a person approves.
No retraining to begin. You improve the agent you already run.
Sikaru is managed continual learning infrastructure for AI agents. You connect an agent you already run or build a new one, and Sikaru runs it and improves it from real work. It manages execution, memory, tools, and evaluations; you guide and approve improvements. Model training is in enterprise beta.
Continual learning means improving a deployed agent from its own production work and user feedback. Sikaru turns misses, user corrections, feedback, and failures into the change that makes the next run better.
Managed agents have their live behavior, safety checks, and improvement handled by Sikaru. You define what good looks like, or let Sikaru infer it, and Sikaru keeps the agent running and improving.
Sikaru catches a signal such as a miss or a correction, finds the pattern, makes the smallest change to the layer that caused it, checks the change against your bar, and makes it live after review so the next run is better.
Context and harness improvement is live today: instructions, knowledge, memory, tools, checks, and orchestration. Model training with RL and SFT on your trajectories is in closed beta with enterprise partners.
No. You can start a new managed agent from scratch and it improves from its first runs. Historical examples help but are not required.
Observability shows what happened and evaluation tools score it. Sikaru turns that signal into reviewed improvements across the agent, so future work gets better from past work.
From your production trajectories. When Sikaru finds a pattern of misses, it builds evals and benchmarks from the real runs behind it. You guide and approve them, and power users can upload their own. RL environments built from your runs are in closed beta with enterprise partners.
No. Every change is graded against your bar and moves from draft to staging to production only after a person on your team approves it.
Not yet. Our SOC 2 Type II audit is in progress. Your work stays in your organization, and we do not train foundation models on it.
Bring your agent, your expertise, and a task that matters. Start a repeatable cycle of feedback, experiments, and reviewed improvements. No dataset required.