Research watch

Developments that change the map.

Curated stories prioritize research implications over volume. Official sources and independent reporting remain visibly distinct.

Curated major stories

Selected for structural relevance.

Each record includes the development and a short explanation of why it matters to this research landscape.

reporting

Benchmark · Mercor / evaluation

APEX tests AI on economically valuable professional work

Mercor's APEX program uses expert-authored tasks to examine professional work in areas such as medicine, law, finance, and consulting.

Why it matters

Evaluation is shifting from exam-style questions toward realistic deliverables and domain-specific judgment.

Read at TIME
official

Company development · AfterQuery / agent training

AfterQuery reports research on agent training for Terminal-Bench 2.0

AfterQuery published work connecting expert-curated trajectories and tooling with improved results on an agent-oriented terminal benchmark.

Why it matters

Training-data providers increasingly combine data production with applied research and evaluation design.

Read at AfterQuery
official

Research direction · Fleet / environments

Fleet centers its research thesis on high-fidelity agent environments

Fleet describes simulated worlds and real-world challenges designed to model work for the training and evaluation of AI agents.

Why it matters

Environment design is emerging as a distinct infrastructure layer for agent development.

Read at Fleet

External feed

Ingestion intentionally inactive.

Planned milestone

A future RSS pipeline will be separated from editorial selections, retain publisher links, and fail gracefully when feeds are unavailable.