Collection of Verification Tasks (MOVED, please follow the link)
⏳ WebGL performance monitor with CPU/GPU load.
A powerful Node.js benchmark library
:rocket: A benchmark script for PHP and MySQL (Archived)
Windows 11 post-install pipeline — audit, tweak, harden, customize
Fair database benchmarks framework and datasets
A simple framework for compile-time benchmarks
JMeter gRPC Request load test plugin for gRPC
UI Benchmark
Fast & more realistic evaluation of chat language models. Includes leaderboard.
[COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
Ansible CIS Benchmark Compliance Remediation for UBUNTU20
Make Dota 2 fps great again
Benchmarking computational single cell ATAC-seq methods
Single Image Deraining: A Comprehensive Benchmark Analysis
:zap: Performance Comparison of Jax-RS implementations and embedded containers
A straightforward JavaScript benchmarking tool and REPL with support for ES modules and libraries.
JSONBench: a Benchmark For Data Analytics On JSON
[ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.
Binary Code Similarity Analysis (BCSA) Benchmark
Evaluating long-term memory of reinforcement learning algorithms
Benchmark for quadratic programming solvers available in Python
A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.
[ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
Automatic download VPR datasets in a standard format
This repository offers a comprehensive library of security policies designed to enhance the security of Kubernetes cluster configurations. The policie...
[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis
A toolbox for benchmarking trustworthiness of multimodal large language models (MultiTrust, NeurIPS 2024 Track Datasets and Benchmarks)
nativejson-benchmark in Rust
Comprehensive CPU frequency performance/power benchmark
This is the official repo for the paper "AMO-Bench: Large Language Models Still Struggle in High School Math Competitions".
Build your own Game-Engine based on the Entity Component System concept in Golang.
Open survey and evidence map for AI agent evolution, self-evolving agents, memory, skills, harnesses, benchmarks, and agent-swarm systems.
Dense Matching Benchmark
Robot Learning Beyond Earth
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
Official Code for What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks (In NeurIPS 2023)
💨 Writing Fast Crystal 😍 -- Collect Common Crystal idioms.
Comparison tools
Ollama based Benchmark with detail I/O token per second. Python with Deepseek R1 example.
Rust libraries and programs focused on succinct data structures
[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive". https://arxiv.org/abs/2404.12...
Benchmarking State-of-the-Art Deep Learning Software Tools
The OpenSSF CVE Benchmark consists of code and metadata for over 200 real life CVEs, as well as tooling to analyze the vulnerable codebases using a va...
Collection of hyperparameter optimization benchmark problems
FBPro Audit Test Automation Package allows you to create compliance reports for your systems. The resulting HTML-reports provide a transparent overvie...
[CVPRW 2022] Delving into High-Quality Synthetic Face Occlusion Segmentation Datasets
The Lodash for GenAI: Real Value + Consistent + Model-Agnostic
A powerful solution as the foundation of your project.