Topic

benchmark

Repositories (1866)

CMExam
CMExam williamliujl Python

A Chinese National Medical Licensing Examination dataset and large languge model benchmarks

97
heurigym
heurigym cornell-zhang Python

Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization (ICLR'26)

96
HLCE
HLCE Humanity-s-Last-Code-Exam Python

(EMNLP 2025 Findings) Source Evaluation scripts for Humanity's Last Code Exam

96
physical-ai-bench
physical-ai-bench SHI-Labs Python

[CVPR 2026 Oral] PAI-Bench: A Comprehensive Benchmark for Physical AI

96
sugar-crepe
sugar-crepe RAIVNLab Python

[NeurIPS 2023] A faithful benchmark for vision-language compositionality

96
agentic-vbench
agentic-vbench PhiloLabs Python

AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

96
MMC
MMC FuxiaoLiu Python

[NAACL 2024] MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning

95
BIRL
BIRL Borda Python

BIRL: Benchmark on Image Registration methods with Landmark validations

95
best
best salesforce TypeScript

:trophy: Delightful Benchmarking & Performance Testing

95
LaBench
LaBench microsoft Go

Latency Benchmarking tool

94
Uniaa
Uniaa KlingAIResearch Python

Unified Multi-modal IAA Baseline and Benchmark

94
physics-benchmarking-neurips2021
physics-benchmarking-neurips2021 cogtoolslab Jupyter Notebook

Repo for "Physion: Evaluating Physical Prediction from Vision in Humans and Machines", presented at NeurIPS 2021 (Datasets & Benchmarks track)

94
drone-tracking
drone-tracking flyers MATLAB

DTB70 -- A Drone Tracking Benchmark

94
gem_bench
gem_bench galtzo-floss Ruby

🪑 Benchmark different versions of same or similar gems & Static Gemfile and installed gem library source code analysis

94
Pano3D
Pano3D VCL3D Python

Code and models for "Pano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth Estimation", OmniCV Workshop @ CVPR21.

94
javascript-serialization-benchmark
javascript-serialization-benchmark adelost TypeScript

Comparison and benchmark of JavaScript serialization libraries (Protocol Buffer, Avro, BSON, etc.)

94
foundational_fsod
foundational_fsod anishmadan23 Python

This repository contains the implementation for the paper "Revisiting Few Shot Object Detection with Vision-Language Models"

93
NexBench
NexBench Nexis-AI TypeScript

NEXBENCH measures what an autonomous agent can actually do on-chain — execute transactions, route swaps, bridge funds, manage DeFi positions, research...

93
PhysBench
PhysBench physical-superintelligence-lab Python

[ICLR 2025] Official implementation and benchmark evaluation repository of <PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical...

93
RSI-Exam
RSI-Exam aiming-lab Python

RSI-Exam: Measuring Recursive Self-Improvement on Long-Horizon, Executable Research Tasks · Discord https://discord.gg/FBzhepyE

93
krun
krun softdevteam Python

High fidelity benchmark runner

92
pibench
pibench sfu-dis C++

Benchmarking framework for index structures on persistent memory

92
serverreview-benchmark
serverreview-benchmark sayem314 Shell

Serverreview Benchmark Script v3

92
HArray
HArray Bazist C++

Fastest Trie structure (Linux & Windows)

92
PointCloudMatters
PointCloudMatters HaoyiZhu Python

[NeurIPS 2024 D&B] Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning

92
pepe
pepe omarmhaimdat Rust

HTTP Load Generator

92
Microbenchmarks
Microbenchmarks JuliaLang C

Microbenchmarks comparing the Julia Programming language with other languages

92
dnstrace
dnstrace redsift Go

Command-line DNS benchmark

92
SUES-200-Benchmark
SUES-200-Benchmark Reza-Zhu Python

SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite

91
MemoryBench
MemoryBench THUIR Python

Code for MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems

91
quantprobe
quantprobe FedericoTs Python

Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your...

91
FeatureBench
FeatureBench LiberCoders Python

[ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"

91
FactCHD
FactCHD zjunlp Python

[IJCAI 2024] FactCHD: Benchmarking Fact-Conflicting Hallucination Detection

91
vizb
vizb goptics Go

A Consistent Data Visualizer for Agents and Humans

91
swt-bench
swt-bench logic-star-ai Python

[NeurIPS 2024] Evaluation harness for SWT-Bench, a benchmark for evaluating LLM repository-level test-generation

90
LOVA3
LOVA3 showlab Python

(NeurIPS 2024) Official PyTorch implementation of LOVA3

90
php-benchmarks
php-benchmarks EFTEC PHP

It is a collection of php benchmarks

90
enterprise-h2ogpte
enterprise-h2ogpte h2oai Python

Client Code Examples, Use Cases and Benchmarks for Enterprise h2oGPTe RAG-Based GenAI Platform

90
vllm-safety-benchmark
vllm-safety-benchmark UCSC-VLAA Python

[ECCV 2024] Official PyTorch Implementation of "How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs"

90
karma-benchmark
karma-benchmark JamieMason TypeScript

A Karma plugin to run Benchmark.js over multiple browsers with CI compatible output.

90
MultiMedEval
MultiMedEval corentin-ryr Python

A Python tool to evaluate the performance of VLM on the medical domain.

90
umesimd
umesimd edanor C++

UME::SIMD A library for explicit simd vectorization.

90
BioDSA
BioDSA RyanWangZf Python

BioDSA: Framework for Vibe Prototyping of AI Agents for Biomedicine

90
step_game
step_game lechmazur

Multi-Agent Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure. A multi-player “step-race” that challenges LLMs to engage i...

90
fast
fast adhocore Go

Check your internet speed/bandwidth right from your terminal. Built on Golang using chromedp

90
locust4j
locust4j myzhan Java

Locust4j is a load generator for locust, written in Java.

89
WildScenes
WildScenes csiro-robotics Python

[IJRR2024] The official repository for the WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Natural Environments

89
sharkbench
sharkbench sharkbench Rust

Benchmarking programming languages and web frameworks.

89
arc-skill
arc-skill pbshgthm Python

An agent skill that plays ARC-AGI-3. One rule: say what an action will do before you spend it. Claude Code on Opus 5 finished all 25 public games at 1...

89
rnn_benchmarks
rnn_benchmarks stefbraun Python

RNN benchmarks of pytorch, tensorflow and theano

89