Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
Time-Series Anomaly Detection | Algorithms + Datasets + Tutorials
Open-Source Framework for Development, Simulation and Benchmarking of Behavior Planning Algorithms for Autonomous Driving
CARLA: A Python Library to Benchmark Algorithmic Recourse and Counterfactual Explanation Algorithms
Research infra for creating RL environments, post-training, and evals
Benchmarks of common ECS (Entity-Component-System)-Frameworks in C++ (or C)
A multi-player tournament benchmark that tests LLMs in social reasoning, strategy, and deception. Players engage in public and private conversations,...
BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
Official implementation for WorldScore: A Unified Evaluation Benchmark for World Generation
High-precision, one-shot and consistent benchmarking framework/harness for Rust. All Valgrind tools at your fingertips.
Benchmarking web frameworks written in rust with rewrk tool.
A simple command line tool to interact with hundreds of servers around the world.
[ITS'21] Human Trajectory Forecasting in Crowds: A Deep Learning Perspective
Official repository of the paper "HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment".
Distributed database benchmark tester
Awesome diffusion Video-to-Video (V2V). A collection of paper on diffusion model-based video editing, aka. video-to-video (V2V) translation. And a vid...
TurboRLE-Fastest Run Length Encoding
Ultra-fast websocket client and server for asyncio
A large-scale benchmark for machine learning methods in fluid dynamics
🔥[NeurIPS'25] DeepFund: Pilot for Your Next Fund Investment
🔥 Synthetic and real-world 2d/3d dataset for semantic and instance segmentation (BMVC 2022 Oral)
BenchExec: A Framework for Reliable Benchmarking and Resource Measurement
Minecraft-style voxel benchmark for comparing AI models (Arena + Sandbox)
JMH benchmark of Java object-to-object mapping frameworks
Pantheon of Congestion Control
Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development.
Automated STIG Benchmark Compliance Remediation for RHEL 7 with Ansible
A distributed storage benchmark for file systems, object stores & block devices with support for GPUs
[TPAMI 2026] Large-Scale 3D Medical Image Pre-training with Geometric Context Priors
TensorFlow Metal Backend on Apple Silicon Experiments (just for fun)
comparing the performance of different template engines
[ICML 2023] Data and code release for the paper "DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation".
Benchmarking framework for protein representation learning. Includes a large number of pre-training and downstream task datasets, models and training/...
Synthetic fraud graph generator for benchmarking graph-based fraud detection models in financial services.
Dataset and code for the paper "First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations", CVPR 2018.
State-of-the-art methods on monocular 3D pose estimation / 3D mesh recovery
PostgreSQL Benchmarking Toolkit
A super simple tool to benchmark GraphQL queries
VideoGen-Eval: Agent-based System for Video Generation Evaluation
Unified World Model Inference & Evaluation Infrastructure
LLM 并发性能测试工具,支持自动化压力测试和性能报告生成。
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
⏱️ single header benchmark framework for C and C++
🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
A micro Vulkan compute pipeline and a collection of benchmarking compute shaders
Benchmarking Agentic LLM and VLM Reasoning On Games
CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
App Servers benchmarked for: Ruby, Python, JavaScript, Dart, Elixir, Java, Crystal, Nim, GO, Rust