Skip to main content
U.S. flag

An official website of the United States government

Demo Codebase for agentic-research-measurement-probes

Published by National Institute of Standards and Technology | National Institute of Standards and Technology | Catalog Last Checked: September 02, 2026 at 07:19 PM | Dataset Last Updated: March 26, 2026
An agentic AI measurement tool for deep research over local corpora of PDF and Markdown documents. Given a research question, it orchestrates a programmatic AI pipeline to exhaustively evaluate, synthesize, and verify information from your documents, producing a Markdown report with inline footnote citations -- then automatically measures the quality of every citation using LM-judge measurement probes. The probes are a first-class feature, not an afterthought. After each report section is written, three mutually exclusive probe evaluators run automatically, scoring every citation along distinct quality dimensions: faithfulness (does the source support the claim?), completeness (is the source's full message represented without cherry-picking?), and sufficiency (does the source carry the evidentiary burden the claim requires, or does the author overreach?). Probe results are stored alongside the report as a structured audit trail, enabling quantitative measurement of AI-generated research quality.

Resources

1 resource available

  • agentic-research-measurement-probes Github Repository

    FILE

Find Related Datasets

Search by Tags

Click any tag below to search for similar datasets

data.gov

An official website of the GSA's Technology Transformation Services

Looking for U.S. government information and services?
Visit USA.gov