Synthetic Sciences, a startup from the Y Combinator W26 batch, released its open-source scientific agent OpenScience from beta on September 27, 2026, and launched it on Product Hunt. The update brought a redesigned workspace, the Ace model service, and Autoresearch mode, where the agent independently runs a series of experiments and improves a specified metric, logging progress in study.md, ideas.md, results.tsv, and lessons.md files. According to published traces, on Terminal-Bench Science, the GPT-6 Astra and GPT-6 Sol combination solved 53 out of 70 tasks (75.7%) compared to 68.1% for Codex with the same lead model. The core is free and open under Apache 2.0, works with your own keys and local models, and monetization has been moved to model payments through the Ace service.

image
image
image

What happened

Synthetic Sciences moved OpenScience from beta to public release on September 27, 2026, and simultaneously launched its Product Hunt page. Users were shown a redesigned workspace and the paid Ace model service with over 30 models, including GPT-6 Astra and Claude Opus 5.5. The central novelty is Autoresearch mode: the agent starts with a baseline run, then iterates through ideas in order of expected utility, deciding at each step whether to keep or roll back the result. After 4 runs without progress, the agent switches to a different type of idea, and every 6 runs it revises its approach. Along with the release, benchmarks were published: on Terminal-Bench Science (70 tasks, 8 hours, one attempt per task), OpenScience with GPT-6 Astra as the lead and GPT-6 Sol as executors solved 53 out of 70 tasks (75.7%) compared to 68.1% for Codex with the same lead model, and on 50 BiomniBench-DA tasks it showed 82.2 compared to 81.04 for the compared configuration. Traces of all these runs are published in a separate benchmarks-openscience repository.

Context

The key detail of the Terminal-Bench Science comparison is that both configurations use the same lead model, GPT-6 Astra, so the gap should be correctly interpreted as a win for the harness and orchestration, not a leap in model capabilities: this is a comparison of scaffolds, not models. Autoresearch mechanically is a managed hill-climbing on a target metric with a budget: the approach inherits AutoML and hyperparameter tuning practices, but instead of a parameter grid, the agent iterates through research ideas in order of expected utility, with stagnation heuristics and stopping by run or time limit. The project consciously positions itself against closed scientific agents: Apache 2.0 license, BYOK own keys, local models, executors on SSH/Slurm/PBS/Modal, and file-based decision logs. The business structure is accordingly: the core is free and open, and payment is tied to model consumption through the Ace service, making entry into autonomous experiments almost free.

Why this matters for the industry

The race for agentic products is shifting from code writing to science, and the first notable effect is methodological: publishing the trace of each benchmark run moves 'verifiability' from a slogan to an artifact that can be downloaded and rechecked step by step, setting a reporting standard where a number cannot be cited without the attached artifact. A free open tool with your own keys puts pressure on paid competitors and lowers the barrier to entry for autonomous experiments to installing an npm package, while the engineering maturity of the framework — executors via SSH/Slurm/PBS/Modal, a metrics tracker with a wandb shim — allows it to be integrated into existing infrastructure, not just shown in demos. At the same time, for Synthetic Sciences itself, the key question of the month is converting installations to model payments through Ace: with almost free entry, unit economics decide everything. If independent groups confirm the results based on the traces, OpenScience could become a standard open baseline for scientific agents, and Autoresearch techniques — budget stop conditions, 'keep/rollback' decisions, idea type change heuristics — will start to be replicated by competing open and commercial products.

Why this matters for users

The tool can already be installed for free: via npm install -g @synsci/openscience or as a desktop application for macOS, Windows, and Linux via the openscience.sh/download link. Then you connect your own Anthropic, OpenAI, or Google keys or a local model, and you can launch your first autonomous experiment with a formulation like 'minimize val_loss, stop after 20 runs or 8 hours' with a metrics panel and familiar tracking via a wandb shim. The cost of a pilot is reduced to your API expenses; the optional Ace service expands the choice to over 30 models, but its prices are not disclosed. Autoresearch logs remain regular files in your project, so a session can be continued, rolled back, or manually analyzed. The claimed benchmark results should not be taken at face value: run traces are in the benchmarks-openscience repository and allow you to go through each step yourself, and a reasonable first run is a limited task on non-critical data with an explicit budget ceiling.

What is still unknown / limitations

All key figures are self-reported by the project: the task set, methodology, and interpretation are set by the vendor itself, and although the traces allow step-by-step rechecking of runs, there is no independent confirmation yet. The statistical base is weak: 82.2 vs 81.04 on 50 BiomniBench-DA tasks is a difference of half a task, and 75.7% vs 68.1% on Terminal-Bench Science is about 5 tasks out of 70 with one attempt per task; without repeated runs and confidence intervals, such gaps cannot be considered established. Ace service prices are not disclosed, nor is data on converting installations to paying Ace customers. The key test for the near term is comparing Codex and Claude Code with the same models and equal budget, as well as reproducing the traces by independent groups. In the long term, the question remains open as to whether autonomous cycles will become a tool for real scientific discoveries: this depends, among other things, on the risk of the agent itself overfitting the target metric.

Sources

Author

Look at AI, editorial team