Research software / Working private prototype

AI agents for
traceable research.

Coordinate coding, evidence gathering, and independent review across long-running projects. SciDocket keeps tasks and research decisions connected to their source files.

For computational labs and R&D teams

SCIDOCKET / WORKFLOWILLUSTRATIVE EXAMPLE
EXAMPLE ASSIGNMENT

Validate a model against its source data.

RESEARCHER / TASK BRIEFSCOPE DEFINED

Define the question and the boundaries.

Specify the source data, fixed assumptions, expected outputs, and checks required before the result can be accepted.

Task brief + source references
Assumptions + acceptance checks

Select a stage to explore. An illustration of the prototype workflow, not a live run.

Prototype integrations
and interfaces
GitHubClaude CodeCodexSlackNotion
01 / The product

Useful work.
A record you can inspect.

Research needs more than a generated answer. SciDocket connects each assignment to its evidence, code changes, checks, and review.

01 / TASKS WITH CONTEXT

One task. A bounded scope.

Give an agent a research question, source material, and an explicit stopping condition. Dependencies determine which work can start.

ObjectiveCheck a model against historical observations.
BoundaryPreserve the agreed model assumptions.
DeliverableReproducible code, comparison, and limitations.
02 / INDEPENDENT REVIEW

A second, separate assessment.

A reviewer checks the result against the task and project requirements. Requested changes return to the same recorded work.

EXAMPLE REVIEW COMMENT

The comparison is reproducible. Explain the treatment of missing observations before this result is adopted.

Execution → Review → Revision or acceptance

Resume from recorded work.

Persistent task state, retained workspaces, and bounded retries support recovery when a worker or session stops.

Keep scientific decisions explicit.

Researchers set project objectives and assumptions. Material changes outside the assigned scope require a human decision.

02 / How it works

From a research question
to a reviewable result.

A persistent workflow connects the tools. GitHub keeps the task, source files, and review history together.

01

Define the work

Record the question, constraints, dependencies, and evaluation criteria in a GitHub task.

OUTPUT / TASK BRIEF
02

Assign an agent

Route eligible work to a configured coding or research backend. Local workers use separate workspaces.

OUTPUT / CODE + EVIDENCE
03

Check the result

Run the relevant tests and route work for independent review before adoption under project rules.

OUTPUT / REVIEW RECORD
04

Retain the findings

Preserve the changes, decisions, and limitations. Approved summaries can also be published to Notion.

OUTPUT / RESEARCH HISTORY
Researcher-defined objectives. Task-specific checks. Traceable handoffs.A negative or inconclusive result can be valid completed work.
03 / Research workflows

Built for work that needs checking.

Start with one bounded workflow in an existing research project. These examples describe intended uses, rather than customer results.

Check what the evidence actually supports.

Assign a focused source review. Record definitions, inclusion decisions, provenance, and unresolved gaps in files that another researcher can inspect.

ILLUSTRATIVE ASSIGNMENT / NO CUSTOMER DATA

REVIEWABLE OUTPUTS

Source inventory and provenance
Explicit inclusion and exclusion decisions
Evidence summary with unresolved questions
04 / Claude in the workflow

A concrete role
for Claude.

Research tasks often combine repository context, code changes, tests, and written interpretation. The Research OS prototype supports Claude Code as an execution and review backend within that process.

Task brief→Claude Code→Review

SciDocket is model-independent. Backend availability depends on configuration and access. Claude integration does not imply Anthropic sponsorship.

Implemented Research OS prototype

Claude Code support

Bounded repository tasks, structured handoffs, and separate execution and review phases. The host checks returned changes before they enter the GitHub workflow.

Planned evaluation Not yet shipped

Claude API integration

Evaluate programmatic execution and review with explicit budgets and task-specific tests. Measure reliability and researcher effort before extending access to more teams.

05 / What exists today

Built from an existing research system.

SciDocket is being developed from Research OS, a private, GitHub-based research orchestrator. The current implementation includes a local runtime, worker interfaces, review handling, and operational views.

Current stage: working private prototype.
Repeatable team onboarding, external pilot evaluation, and a Claude API path are next-stage work. A public hosted product is not yet available.

GitHub-native research records

Issues, dependencies, branches, and pull requests retain task context and review history.

Persistent local orchestration

Task claims, retries, execution receipts, and restart recovery support long-running work.

Separate worker and host responsibilities

The host manages Git changes and workflow transitions. Local workers return bounded outputs for validation.

Slack, dashboard, and curated knowledge

Slack provides operator input; a read-only dashboard exposes status; an optional Notion integration publishes retained summaries.

06 / Founder

Built by a researcher
for research work.

SciDocket addresses a practical coordination problem: using AI across projects while retaining the assumptions, source material, and checks needed to evaluate the results.

The goal is to make this workflow usable by other computational research groups and small R&D teams.

Baptiste Andrieu, PhD

FOUNDER / RESEARCHER

A researcher at the University of Cambridge, working on material supply chains and computational modelling. His background includes stock-and-flow models and evidence-intensive environmental research.

SciDocket is an independent project. The founder's university affiliation describes his background and does not imply institutional endorsement.

Before a pilot

A few practical details.

Can my team use SciDocket today?

The software is a working private prototype. We are seeking conversations with computational labs and R&D teams about a bounded pilot. Team onboarding and broader access are still being developed; there is no public self-service signup.

Does SciDocket make scientific decisions autonomously?

Agents can propose methods, execute tasks, and challenge assumptions within their assignment. The project defines which results can be adopted. Material changes of scientific direction outside the authorized scope require a human decision. Review does not guarantee scientific correctness.

Where do my code and data go?

The prototype runs from a local host and records work in GitHub. AI providers may receive the content needed for assigned tasks, according to the configured backend. Local orchestration does not mean that all processing stays on your machine. Confidential data and institutional requirements must be assessed before a pilot.

Does it require Claude for every task?

No. The prototype supports configured coding backends such as Claude Code and Codex, alongside explicit handoffs to browser-based research workers. The proposed Claude API integration is future work. Models and accounts do not replace the task record or the review process.

What would an early pilot evaluate?

A useful first pilot would test one research workflow: whether outputs can be reproduced, whether another researcher can inspect the evidence, how often intervention is needed, and how failures are recovered. These are evaluation objectives, not performance results already established.

Let's start with one workflow

Make your next research task inspectable.

Tell us what your team is working on, which tools you use, and where coordination becomes difficult.

Discuss a pilot Working private prototype.
Pilot scope and availability discussed individually.