Benchmark for some popular PHP Dependency Injection Containers.
Store data created during your `pytest` tests execution, and retrieve it at the end of the session, e.g. for applicative benchmarking purposes.
A low-latency LRU approximation cache in C++ using CLOCK second-chance algorithm. Multi level cache too. Up to 2.5 billion lookups per second.
a http server benchmark tool written in rust 🦀
🔮 Obtain the power of touchless interaction with display screens
Blazing fast AndroidX DocumentFile alternative for Android SAF (scoped storage). Up to ~14x faster on large directories.
Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
IPC benchmark on Linux
Cache benchmark for Golang
[ICLR 2023 spotlight] MEDFAIR: Benchmarking Fairness for Medical Imaging
A list of LLM benchmark frameworks.
Official code for "CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges"
StepShield: When, Not Whether to Intervene on Rogue Agents — NeurIPS 2026 benchmark for temporal evaluation of AI agent guardrails (9,429 trajectories...
The source code and data link of paper "MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare".
Benchmark local LLM serving by replaying real agent trajectories
Write benchmarks without the hassle.
Benchmark scripts for TVM
(ECCV2026) RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing
Modern C++ benchmarking
RPC Benchmark of gRPC, Aeron and KryoNet
A Python and MATLAB implementation of mathematical test functions for benchmarking optimization algorithms.
Fastest Histogram Construction
Official repository for KoMT-Bench built by LG AI Research
Vector Index Benchmark for Embeddings (VIBE) is an extensible benchmark for approximate nearest neighbor search methods, or vector indexes, using mod...
A project about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, es...
[IJCAI-2021] Contrastive Model Inversion for Data-Free Knowledge Distillation
Thematic Generalization Benchmark: measures how effectively various LLMs can infer a narrow or specific "theme" (category/rule) from a small set of ex...
Code and data for "ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM" (NeurIPS 2024 Track Datasets and Benchmarks)
Benchmarking Rust key-value storage engines
Web Components benchmark for a various Web Components technologies
The benchmark of ncnn that is a high-performance neural network inference framework optimized for the mobile platform
Benchmarks: write in Scala or JS, run in your browser. Live demo:
RNN-based Codon Optimization Tool. Publication: https://doi.org/10.1186/s12859-023-05246-8
[AAAI 2021] (oral) Progressive One-shot Human Parsing, [TPAMI 2023] End-to-end One-shot Human Parsing
[ICML 2026] GenExam: A Multidisciplinary Text-to-Image Exam
Open benchmark on observability tasks built on Harbor
Simulator + benchmark suite for Micro Aerial Vehicle design.
MoleculeNet benchmark dataset & MolMapNet dataset
When storing a value in a Go interface allocates memory on the heap.
[NeurIPS 2024] A task generation and model evaluation system for multimodal language models.
Modern JavaScript benchmarking tool.
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
A comparison of many lossless image compression formats.
Agent eval on one running base. Swap the agent under test with plugins; run the same dataset anywhere.
A Deep Journey into Super-resolution: A Survey, ACM Computing Surveys
A Survey and Benchmark of QUIC
Large-scale uncertainty benchmark in deep learning.
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
Audio-Oscar is a multi-agent framework for generating long-form, controllable audio from complex audio scene descriptions.
Reproducible on-device LLM benchmarks for Apple Silicon (iPhone 17 Pro, M4 Max): Apple Core AI, MLX, llama.cpp, LiteRT-LM and Core ML on the same mode...