Research Organization

Simone Systems Research

Independent AI Systems Research

Research on how AI systems coordinate, verify results, allocate compute, and improve reliably.

Current work focuses on agent orchestration, evaluation, verification, compute economics, and adaptive AI systems.

Research Focus

Core questions guiding our systems development and empirical investigations.

Systems & Control

Agent Orchestration

How should heterogeneous models, tools, and agents divide work and coordinate reliably?

Reliability & Measurement

Evaluation & Verification

How do we distinguish genuine improvement from false progress, benchmark noise, and brittle behavior?

Inference & Budgeting

Compute Economics

When does additional inference improve verified outcomes enough to justify its cost?

Dynamic Learning

Adaptive AI Systems

How can systems learn from measured failures and improve their own workflows without losing control or reproducibility?

Research Projects

Prototypes, evaluation suites, and experimental systems implementations.

GitHub Profile →

The Fleet

Active

How can heterogeneous agent topologies coordinate deterministically under strict latency and tool constraints?

Evidence-Gated Evaluation

Active

What evaluation methodologies prevent benchmark overfitting and reliably detect subtle regression in multi-turn agents?

Cross-Model Verification

Experimental

Can asymmetric verifier models provide formal safety bounds for black-box generative agents?

Pokémon Policy Evolution

Active

Empirical study on measured state tracking, environmental feedback, and policy adaptation in discrete environments.

Research Principles

Standard operating commitments applied across all experiments, evaluations, and published artifacts.

01

Evidence before promotion

Improvements should survive measurement, not merely look plausible.

02

Independent verification

Positive results should be reproduced outside the context that generated them.

03

Compute must earn its cost

Expensive inference should be reserved for decisions where it materially improves verified outcomes.

04

Negative results are retained

Failed hypotheses are useful evidence and should not disappear from the record.

05

Artifacts matter

Claims should refer to the exact code, model, configuration, and environment actually evaluated.

Research Notes

Working papers, benchmark reports, and technical reflections.

All Notes →