Logs performance benchmark repo: Comparing Elastic, Loki and SigNoz
First LLM Studio: local-first LLM studio for Apple Silicon with MLX runtimes, Compare Lab, benchmark ops, replay, and runtime telemetry.
:boom: Performance-focused HTTP load testing tool written in Go
Simple DNS bench util that supports encrypted protocols.
Learned Sort: a model-enhanced sorting algorithm
基于Python Tornado的高性能http性能测试工具。Java Netty版: https://github.com/junneyang/http-benchmark-netty 。
The benchmark to compare performance of PHP ORM solutions.
Data race benchmark suite for evaluating OpenMP correctness tools aimed to detect data races.
🚀 Spiko is a fast, Rust-based load testing tool with a beautiful TUI for real-time insights.
A benchmark dataset collection for bird sound classification
🚀 A comprehensive performance comparison benchmark between different .NET collections.
Framework for benchmarking fully-managed vector databases
[NeurIPS 2024] Terra: A Multimodal Spatio-Temporal Dataset Spanning the Earth
Automated Benchmarking System for Vitess
A benchmark framework based on Golang
[ACL 2025 🔥] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
Compare performance of macOS browsers based on Speedometer 3.1
A digital representation of Sikh Bani and other Panthic texts with a public logbook of sangat-sourced corrections.
Agent Memory Benchmark
Benchmark your 3DS battery
LeakDB (Leakage Diagnosis Benchmark) is a realistic leakage dataset for water distribution networks. The dataset is comprised of a large number of art...
Blazing fast AndroidX DocumentFile alternative for Android SAF (scoped storage). Up to ~14x faster on large directories.
A benchmark for testing whether coding agents can resolve engineering tasks in scientific software
Safe Multi-Agent MuJoCo benchmark for safe multi-agent reinforcement learning research.
The Zebrafish Activity Prediction Benchmark measures progress on the problem of predicting cellular-resolution neural activity throughout an entire ve...
An open collaborative repository for reproducible specifications of HPC benchmarks and cross site benchmarking environments
🔮 Obtain the power of touchless interaction with display screens
Store data created during your `pytest` tests execution, and retrieve it at the end of the session, e.g. for applicative benchmarking purposes.
A low-latency LRU approximation cache in C++ using CLOCK second-chance algorithm. Multi level cache too. Up to 2.5 billion lookups per second.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
Benchmark for some popular PHP Dependency Injection Containers.
a http server benchmark tool written in rust 🦀
[ICLR 2026] IVEBench - Benchmark for Instruction-Guided Video Editing
StepShield: When, Not Whether to Intervene on Rogue Agents — NeurIPS 2026 benchmark for temporal evaluation of AI agent guardrails (9,429 trajectories...
Cache benchmark for Golang
IPC benchmark on Linux
Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
A list of LLM benchmark frameworks.
Official code for "CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges"
[ICLR 2023 spotlight] MEDFAIR: Benchmarking Fairness for Medical Imaging
Benchmark scripts for TVM
First-of-its-kind AI benchmark for evaluating the protection capabilities of large language model (LLM) guard systems (guardrails and safeguards)
Write benchmarks without the hassle.
RPC Benchmark of gRPC, Aeron and KryoNet
Official repository for KoMT-Bench built by LG AI Research
[IJCAI-2021] Contrastive Model Inversion for Data-Free Knowledge Distillation
Modern C++ benchmarking
A Python and MATLAB implementation of mathematical test functions for benchmarking optimization algorithms.
Fastest Histogram Construction
Code and data for "ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM" (NeurIPS 2024 Track Datasets and Benchmarks)