Topic

benchmark

Repositories (1827)

Mind2Web-2
Mind2Web-2 OSU-NLP-Group Python

[NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge

112
holisticai
holisticai holistic-ai Jupyter Notebook

This is an open-source tool to assess and improve the trustworthiness of AI systems.

112
AIDABench
AIDABench MichaelYang-lyx Python

Code for paper AIDABench: AI Data Analytics Benchmark.

111
text2image-benchmark
text2image-benchmark boomb0om Jupyter Notebook

Benchmark for generative image models

111
OmniBenchmark
OmniBenchmark ZhangYuanhan-AI Python

[ECCV2022] New benchmark for evaluating pre-trained model; New supervised contrastive learning framework.

110
benchyou
benchyou xelabs Go

benchyou is a benchmark tool for MySQL, real-time monitoring TPS and vmstat/iostat

110
kubernetes-iperf3
kubernetes-iperf3 Pharb Shell

Simple wrapper around iperf3 to measure network bandwidth from all nodes of a Kubernetes cluster

110
mini-nbody
mini-nbody harrism C

A simple gravitational N-body simulation in less than 100 lines of C code, with CUDA optimizations.

109
VisualNews-Repository
VisualNews-Repository FuxiaoLiu Jupyter Notebook

[EMNLP'21] Visual News: Benchmark and Challenges in News Image Captioning

109
tsnkit
tsnkit ChuanyuXue Python

A scheduling and benchmark toolkit for Time-Sensitive Networking in Python

109
pytorch-benchmark
pytorch-benchmark LukasHedegaard Python

Easily benchmark PyTorch model FLOPs, latency, throughput, allocated gpu memory and energy consumption

109
build-tools-performance
build-tools-performance rstackjs JavaScript

Benchmarks for bundlers and build tools, including Rspack, Rsbuild, webpack, Vite, Rolldown, esbuild, Parcel, Farm and Utoo.

108
kaggle-dogs-vs-cats-caffe
kaggle-dogs-vs-cats-caffe mrgloom Python

Kaggle dogs vs cats solution in Caffe

107
coder_eval
coder_eval UiPath Python

Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B expe...

107
cwe-bench-java
cwe-bench-java iris-sast Python

A manually vetted dataset for security vulnerability detection in Java projects

107
EvalNE
EvalNE Dru-Mara Python

Source code for EvalNE, a Python library for evaluating Network Embedding methods.

107
peaks-consolidation
peaks-consolidation hkpeaks Go

The Peaks Consolidation is equipped with state-of-the-art algorithms and data structures that support high-performance databending exercises. It speci...

107
MCC5-THU-Gearbox-Benchmark-Datasets
MCC5-THU-Gearbox-Benchmark-Datasets liuzy0708 MATLAB

A benchmark fault diagnosis dataset comprises vibration data collected from a gearbox under variable working conditions with intentionally induced fau...

107
apebench
apebench tum-pbs Python

[Neurips 2024] A benchmark suite for autoregressive neural emulation of PDEs. (≥46 PDEs in 1D, 2D, 3D; Differentiable Physics; Unrolled Training; Roll...

107
video_object_detection_paper
video_object_detection_paper junliang230

update some video object detection papers (视频目标检测论文和代码整理)

106
hash-smith
hash-smith bluuewhale Java

Fast & memory efficient hash tables for Java

106
vault-benchmark
vault-benchmark hashicorp Go

A tool for benchmarking usage of Vault.

106
pplbench
pplbench facebookresearch Python

Evaluation Framework for Probabilistic Programming Languages

105
mlip-arena
mlip-arena atomind-ai Jupyter Notebook

🌟 [NeurIPS '25 Spotlight] Fair and transparent benchmark of machine learning interatomic potentials (MLIPs), beyond basic error metrics https://openr...

105
benchmark-websocket
benchmark-websocket oatpp C++

Websocket Client and Server for benchmarks with Millions of concurrent connections.

104
cec2017-py
cec2017-py tilleyd Python

Python module for CEC 2017 single objective optimization test function suite.

104
solidity-benchmarks
solidity-benchmarks alephao Solidity

Benchmarks of popular contract implementations in solidity

104
ddio-bench
ddio-bench aliireza Makefile

Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit Networks

104
tastylib
tastylib chuyangliu C++

C++ implementations of data structures, algorithms, and system designs.

104
deepmark
deepmark IngestAI PHP

Deepmark AI enables a unique testing environment for language models (LLM) assessment on task-specific metrics and on your own data so your GenAI-powe...

104
YAIB
YAIB rvandewater Python

🧪Yet Another ICU Benchmark: a holistic framework for the standardization of clinical prediction model experiments. Provide custom datasets, cohorts,...

104
PAD
PAD EricLee0224 Python

[NeurlPS 2023] A Dataset and Benchmark for Pose-agnostic Anomaly Detection.

104
zk-Harness
zk-Harness zkCollective Python

Benchmarking framework for general purpose zero-knowledge proofs languages and libraries

103
trajectopy
trajectopy gereon-t Python

Trajectopy - Trajectory Evaluation in Python

103
RealPDEBench
RealPDEBench AI4Science-WestlakeU Python

[🔥ICLR26 Oral] RealPDEBench: A Benchmark for Complex Physical Systems with Paired Real-World and Simulated Data

103
annbench
annbench matsui528 Python

A lightweight benchmark for approximate nearest neighbor search

103
datacenter-speed-tests
datacenter-speed-tests jakejarvis Shell

⚡ Test speed and pings to all DigitalOcean, Linode, AWS, GCP, and Vultr regions

102
playwright-test
playwright-test hugomrdias JavaScript

Run unit tests with several test runners or benchmark inside real browsers with playwright and other Javascript runtimes.

102
ansibench
ansibench nfinit C

A selection of ANSI C benchmarks and programs useful as benchmarks

102
PPM
PPM ZHKKKe

A High-Quality Photograpy Portrait Matting Benchmark

102
AIRTBench-Code
AIRTBench-Code dreadnode Jupyter Notebook

Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

102
XRAutomatedTests
XRAutomatedTests Unity-Technologies C#

XRAutomatedTests is where you can find functional, graphics, performance, and other types of automated tests for your XR Unity development.

102
MMRole
MMRole YanqiDai Python

[ICLR 2025] A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents

101
dm_nevis
dm_nevis google-deepmind Python

NEVIS'22: Benchmarking the next generation of never-ending learners

101
PhysGym
PhysGym principia-ai Python

A benchmark suite for evaluating LLM-based interactive scientific reasoning.

101
mqtt-mock
mqtt-mock daoshenzzg Go

mqtt压测工具。支持subscribe、publish压测方式,支持模拟客户端连接数。

101
GLUE-X
GLUE-X YangLinyi Python

We leverage 14 datasets as OOD test data and conduct evaluations on 8 NLU tasks over 21 popularly used models. Our findings confirm that the OOD accur...

100
Wild-Places
Wild-Places csiro-robotics Python

🏞️ [IEEE ICRA2023] The official repository for paper "Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural Environm...

100
Robust-Gymnasium
Robust-Gymnasium SafeRL-Lab Python

[ICLR 2025] Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning.

100
RWKU
RWKU jinzhuoran Python

RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models. NeurIPS 2024

100