Cista is a simple, high-performance, zero-copy C++ serialization & reflection library.
Benchmarking Knowledge Transfer in Lifelong Robot Learning
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
:zap: Go web framework benchmark
A machine learning toolkit for log parsing [ICSE'19, DSN'16]
An objective comparison of multiple frameworks that allow us to "transform" our web apps to desktop applications.
Tracking Any Point (TAP)
benchmarks for implementation of servers which support 1 million connections
Playing around "Less Slow" coding practices in C++ 20, C, CUDA, PTX, & Assembly, from numerics & SIMD to coroutines, ranges, exception handling, netwo...
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
Efficient Retrieval Augmentation and Generation Framework
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference implementations of MLPerf® training benchmarks
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
Simple, fast, accurate single-header microbenchmarking functionality for C++11/14/17/20
Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.
SkillsBench evaluates how well skills work and how effective agents are at using them.
BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx
Reference implementations of MLPerf® inference benchmarks
The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
:bar_chart: Benchmark multiple object trackers (MOT) in Python
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
pytest fixture for benchmarking code
C++20 μ(micro)/Unit Testing framework
Fast and simple benchmarking for Rust projects
[pip install medmnist] 18x Standardized Datasets for 2D and 3D Biomedical Image Classification
Fast Compiler for C# Expression Trees and the lightweight LightExpression alternative. Diagnostic and code generation tools for the expressions.
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL...
SMAC: The StarCraft Multi-Agent Challenge
Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book "Foundations of Deep Reinforcement Learning".
Latest Advances on System-2 Reasoning
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
jsperf.com v2. https://github.com/h5bp/lazyweb-requests/issues/174
Microbenchmarking app for Swift with nice log-log plots
[NeurIPS '25] Knowledge Graph Generation from Any Text
GitHub Action for continuous benchmarking to keep performance
A better load generator for locust, written in golang.
LongBench v2 and LongBench (ACL 25'&24')
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
A compilation of Linux server benchmarking scripts.
Computational framework for reinforcement learning in traffic control
PDEBench: An Extensive Benchmark for Scientific Machine Learning
Application Performance Optimization Summary
Benchmark for vector databases.
JMLR: OmniSafe is an infrastructural framework for accelerating SafeRL research.
Tracking Benchmark for Correlation Filters
OpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive Learning
🚀 Fast prime number generator
ClickBench: a Benchmark For Analytical Databases
lzbench is an in-memory benchmark of open-source compressors