Build company as agent.

RSIX makes operating a company the training signal. The system acts in a real market, measures what the action was worth, and updates its own policy on the result.

In the past, companies learned through people.

Experience accumulated in individuals and left when they did. We are making the operating itself the training: every decision carries a hypothesis, every action produces a measured outcome, and every verified outcome updates the policy that produced it.

Recursive self-improvement is to us what the recommendation algorithm was to ByteDance — the growth and decision engine the whole system runs on. If recommendation taught a company to understand people and content, RSI teaches it to understand actions and their consequences.

system policy action self-grade acts scores updates Self-grading the signal never leaves system policy action market realized profit acts sells, spends cash updates Market-grounded the world prices the result
Figure 1 — Act, outcome, evaluate, learn, update policy. Both systems run that loop; only one lets the world price the result. Remove the two edges that leave the system and return carrying cash, and improvement is being measured by the thing being improved.

In the age of AI, the decisive competitive variable is evolution speed.

Three stages: actionable, then learnable, then compounding.

Each stage is a precondition for the next. An environment that cannot be acted on produces no signal; actions that cannot be evaluated produce no learning; learning that is not retained produces no advantage.

Stage 01

Agentizationmake the environment actionable

Turn the company from a software system into an environment an AI can operate. In commerce that means the entire chain — demand, supply, and the venue where they meet.

demand user agent intent, unstated need supply product agent fit, durability, use venue platform agent agentic commerce full context full context transaction and fulfillment matches outcome
Figure 2 — Both sides of the market are rebuilt as full-context agents, so intent and product reality meet as structured objects rather than keywords and images. The platform agent closes the transaction, and the outcome returns as signal.

Stage 02

Recursive self-improvementmake actions learnable

The system proposes experiments, executes decisions, evaluates outcomes and updates its own agents, skills and policies — turning every commercial act into a verifiable experiment. Improvement happens at three depths.

evaluator realized profit L3 · architecture which agents exist, and how they connect L2 · procedures the prompts and operating steps each agent follows L1 · policies matching, bidding and pricing models rebuilds rewrites retrains
Figure 3 — One evaluator, three depths. Most systems stop at L1. The loop only becomes recursive when the same evidence that retrains a model is also allowed to rewrite the procedure that produced it, and redraw the graph of agents around it.

Stage 03

Proprietary experiencemake learning compound

Each verified result becomes training data for the next decision. Action produces a real outcome, the outcome is verified, and the verified outcome improves the policy that produced it — a data flywheel no one can buy.

a traditional company capital inventory + marketing revenue the next cycle starts from the same place the asset is consumed RSIX capital experiments real outcomes verified experience priors · policies · causal knowledge verified compounds the asset accumulates
Figure 4 — The same dollar, two destinations. Both companies turn capital into revenue; only one leaves something behind when the cycle closes. The lower row has one node the upper row cannot have without owning the environment — and it is the only node whose value goes up every time the loop runs.

A traditional company converts capital into inventory. We convert capital into intelligence.

Commerce is the proving ground, not the destination.

We extend the loop only after it survives real economic measurement — orders, contribution profit, returns, repeat purchase and fulfillment, not a benchmark score.

Now RSI core evaluator · experiments · policy eCommerce acts profit RSI for eCommerce Next RSI core unchanged commerce social gaming profit RSI for X Long term objective · capital authority · risk boundary RSI core unchanged a business it stands up acts profit Company Agent
Figure 5 — The core never changes. What changes is how many environments it is wired into, and who specifies the objective. Phase two is the claim that the loop, not the domain knowledge, is the transferable asset; phase three is the point where the dashed boxes — the brief and the business — are the only things left for a person to supply.

We are not building agents to do people’s work. We are building companies that evolve.

A research team that owns the environment it studies.

Most work on recursive self-improvement is validated against benchmarks, because a benchmark is what a lab can hold. We hold something rarer: a live economy, with a budget, a supply chain and the authority to act in it. Our results are graded by what the market pays.

The team has trained frontier recommendation, advertising and generative systems at the scale of hundreds of millions of users, and has been accountable for what those systems earned — the combination this problem requires and almost no one has.

Research areas

  • Recursive self-improvement
  • Agent learning and post-training
  • Reward modeling and evaluation
  • Causal inference and experiment design
  • Generative recommendation
  • World models and simulation
  • Multimodal generation
  • Training and online-learning infrastructure

Where the team comes from

  • ByteDance
  • Google
  • Meta
  • Kuaishou
  • Tencent

We are hiring the people who will build the loop.

talent@rsix.ai

Research and engineering: agent learning, evaluation, causal inference, recommendation, infrastructure.

hello@rsix.ai

Partnership, supply and capital.