NoSQL Redis and Memcache traffic generation and benchmarking tool.
TorchBench is a collection of open source benchmarks used to evaluate PyTorch performance.
PHP Framework Benchmark
Python Performance Benchmark Suite
Official Implement of "ADBench: Anomaly Detection Benchmark", NeurIPS 2022.
Model Zoo For OpenCV DNN and Benchmarks.
Airspeed Velocity: A simple Python benchmarking tool with web-based reporting
Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models
Monocular Depth Estimation Toolbox based on MMSegmentation.
BlazeHTTP 是一款简单易用的 WAF 防护效果测试工具。BlazeHTTP stands as a user-friendly WAF protection efficacy evaluation tool.
Various gRPC benchmarks
A High Performance HTTP Server for Ruby
CUDA Kernel Benchmarking Library
VPS benchmark script — based on the popular bench.sh, plus CPU and ioping tests, and dual-stack IPv4 and v6 speedtests by default
Measure Amazon S3's performance from any location.
A PyTorch library for all things Reinforcement Learning (RL) for Combinatorial Optimization (CO)
Python suite to construct benchmark machine learning datasets from the MIMIC-III 💊 clinical database.
AoE (AI on Edge,终端智能,边缘计算) 是一个终端侧AI集成运行时环境 (IRE),帮助开发者提升效率。
🐰 Bencher - Continuous Benchmarking
Performance comparison of .NET IoC containers
Windows, macOS and Android storage (HDD, SSD, RAM) speed testing/performance benchmarking app
C++ Benchmark Authoring Library/Framework
[CBLUE1] 中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
Natural Intelligence is still a pretty good idea.
A benchmark dataset for data-driven weather forecasting
📊 Benchmark Comparison of Packages with Runtime Validation and TypeScript Support
High-performance Distributed Storage
[NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial rel...
S3 benchmarking tool
A dataset of datasets for learning to learn from few examples
Yet another implementation of computer language benchmarks game
"Trust no one, bench everything." - sbt plugin for JMH (Java Microbenchmark Harness)
Internal Safety Collapse: Turning the LLM or an AI Agent into a sensitive data generator.
RobustBench: a standardized adversarial robustness benchmark [NeurIPS 2021 Benchmarks and Datasets Track]
Easily monitor your ThreeJS performances.
HammerDB: The industry standard open-source database benchmark
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted...
golang HTTP stress testing tool, support single and distributed, http/1, http/2 and http/3.
An agent benchmark with tasks in a simulated software company.
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
Evaluation of the CNN design choices performance on ImageNet-2012.
Tasks Assessing Protein Embeddings (TAPE), a set of five biologically relevant semi-supervised learning tasks spread across different domains of prote...
A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.
Another benchmark for some python frameworks
Raw benchmarks on throughput, latency and transfer of Hello World on popular microservices frameworks
A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper
comparing the c ffi (foreign function interface) overhead on various programming languages
Frontier Models playing the board game Diplomacy.