Hey there, friend using a screen reader. I've carefully labeled every button and image on this site. Hope you have a smooth experience. If anything is inconvenient, my email is in the footer — just let me know and I'll fix it. — Store Owner
In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot. The post Project HydraFusion: Frontier quality via multi-model orchestration appeared first on The…
GitHubPublished Sep 5, 2026
Official source · GitHub AI & MLVerified Oct 5, 2026
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictiv…
Microsoft ResearchPublished Aug 21, 2026
Official source · Microsoft ResearchVerified Oct 10, 2026
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared…
Microsoft ResearchPublished Aug 4, 2026
Official source · Microsoft ResearchVerified Oct 6, 2026
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
Microsoft ResearchPublished Aug 13, 2026
Official source · Microsoft ResearchVerified Oct 10, 2026
We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog .
GitHubPublished Oct 5, 2026
Official source · GitHub AI & MLVerified Oct 10, 2026
These are the lessons we learned evaluating LLMs for real-world secret scanning. The post How to evaluate LLMs before production appeared first on The GitHub Blog .
GitHubPublished Aug 26, 2026
Official source · GitHub AI & MLVerified Sep 24, 2026
Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ). To this end, several arena-based leaderboards have established themselves as useful reference points for the community:
Hugging FacePublished Sep 30, 2026
Cross-checked · Hugging Face BlogVerified Oct 10, 2026
AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025 , and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema . This next phase of the collaboration puts that shared infrastructure into practice.
Hugging FacePublished Sep 22, 2026
Cross-checked · Hugging Face BlogVerified Oct 10, 2026
Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.
Hugging FacePublished Sep 2, 2026
Cross-checked · Hugging Face BlogVerified Sep 23, 2026
Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:
Hugging FacePublished Aug 28, 2026
Cross-checked · Hugging Face BlogVerified Sep 22, 2026
One reason is that traditional benchmarks overlook many of the conditions and qualities that make voice systems reliable, natural, contextually appropriate, and effective in practice. That's why we recently introduced held-out sets in Real World VoiceEQ , the Open-ASR Leaderboard , and the Far-field ASR Leaderboard :…
Hugging FacePublished Aug 21, 2026
Cross-checked · Hugging Face BlogVerified Sep 9, 2026