[NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge
This is an open-source tool to assess and improve the trustworthiness of AI systems.
Code for paper AIDABench: AI Data Analytics Benchmark.
Benchmark for generative image models
[ECCV2022] New benchmark for evaluating pre-trained model; New supervised contrastive learning framework.
benchyou is a benchmark tool for MySQL, real-time monitoring TPS and vmstat/iostat
Simple wrapper around iperf3 to measure network bandwidth from all nodes of a Kubernetes cluster
A simple gravitational N-body simulation in less than 100 lines of C code, with CUDA optimizations.
[EMNLP'21] Visual News: Benchmark and Challenges in News Image Captioning
A scheduling and benchmark toolkit for Time-Sensitive Networking in Python
Easily benchmark PyTorch model FLOPs, latency, throughput, allocated gpu memory and energy consumption
Benchmarks for bundlers and build tools, including Rspack, Rsbuild, webpack, Vite, Rolldown, esbuild, Parcel, Farm and Utoo.
Kaggle dogs vs cats solution in Caffe
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B expe...
A manually vetted dataset for security vulnerability detection in Java projects
Source code for EvalNE, a Python library for evaluating Network Embedding methods.
The Peaks Consolidation is equipped with state-of-the-art algorithms and data structures that support high-performance databending exercises. It speci...
A benchmark fault diagnosis dataset comprises vibration data collected from a gearbox under variable working conditions with intentionally induced fau...
[Neurips 2024] A benchmark suite for autoregressive neural emulation of PDEs. (≥46 PDEs in 1D, 2D, 3D; Differentiable Physics; Unrolled Training; Roll...
update some video object detection papers (视频目标检测论文和代码整理)
Fast & memory efficient hash tables for Java
A tool for benchmarking usage of Vault.
Evaluation Framework for Probabilistic Programming Languages
🌟 [NeurIPS '25 Spotlight] Fair and transparent benchmark of machine learning interatomic potentials (MLIPs), beyond basic error metrics https://openr...
Websocket Client and Server for benchmarks with Millions of concurrent connections.
Python module for CEC 2017 single objective optimization test function suite.
Benchmarks of popular contract implementations in solidity
Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit Networks
C++ implementations of data structures, algorithms, and system designs.
Deepmark AI enables a unique testing environment for language models (LLM) assessment on task-specific metrics and on your own data so your GenAI-powe...
🧪Yet Another ICU Benchmark: a holistic framework for the standardization of clinical prediction model experiments. Provide custom datasets, cohorts,...
[NeurlPS 2023] A Dataset and Benchmark for Pose-agnostic Anomaly Detection.
Benchmarking framework for general purpose zero-knowledge proofs languages and libraries
Trajectopy - Trajectory Evaluation in Python
[🔥ICLR26 Oral] RealPDEBench: A Benchmark for Complex Physical Systems with Paired Real-World and Simulated Data
A lightweight benchmark for approximate nearest neighbor search
⚡ Test speed and pings to all DigitalOcean, Linode, AWS, GCP, and Vultr regions
Run unit tests with several test runners or benchmark inside real browsers with playwright and other Javascript runtimes.
A selection of ANSI C benchmarks and programs useful as benchmarks
A High-Quality Photograpy Portrait Matting Benchmark
Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
XRAutomatedTests is where you can find functional, graphics, performance, and other types of automated tests for your XR Unity development.
[ICLR 2025] A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents
NEVIS'22: Benchmarking the next generation of never-ending learners
A benchmark suite for evaluating LLM-based interactive scientific reasoning.
mqtt压测工具。支持subscribe、publish压测方式,支持模拟客户端连接数。
We leverage 14 datasets as OOD test data and conduct evaluations on 8 NLU tasks over 21 popularly used models. Our findings confirm that the OOD accur...
🏞️ [IEEE ICRA2023] The official repository for paper "Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural Environm...
[ICLR 2025] Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning.
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models. NeurIPS 2024