One-shot LLM eval cases by NAGI STUDIO - same prompt, different agents (model + harness), runnable artifacts side by side.
Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Self-contained, Dockerized offensive security challenges for evaluating AI-powered penetration testing agents. Covers modern tech stacks (Node.js, Pyt...
Directly calling functional components instead of mounting them is faster.
Open Source AI Benchmarking toolkit for benchmarking speech to text services
Smart benchmarking of pull requests with statistical confidence
Dataset and Evaluation Scripts for Obstacle Detection via Semantic Segmentation in a Marine Environment
Distributed S3 benchmarking tool - Replacement of Cosbench
Wide NoSQL benchmark for RocksDB, LevelDB, Redis, WiredTiger and MongoDB extending the Yahoo Cloud Serving Benchmark
A benchmark for Salient Object Detection (SOD).
利用 Fetch API 向目标网站发送频繁请求,模拟按下 F5 刷新的效果,以测试服务器的资源限制。
🌈 Visualizes your BenchmarkDotNet benchmarks to Colorful images and Feature-rich HTML (and maybe powerful charts in the future!)
[SCI-FM@ICLR 2025] Specialized LLMs capable of handling various diabetes tasks
Official repository for the paper "ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming"
Can LLMs beat classical HPO? A benchmark comparing classical, LLM-based, and hybrid methods on Karpathy's autoresearch.
LLM benchmark and leaderboard for narrator-bias sycophancy, opposite-narrator contradictions, and judgment consistency.
Lightweight Python tool using Optuna for tuning llama.cpp flags: towards optimal tok/s for your machine
掘金的小册《Android 进阶:基于 Kotlin 的 Android App 开发实践》中的相关的例子
Audio performance benchmark for jitter, theoretical latency, etc.
Salient objects in clutter, TPAMI, 2022
[ICLR 2025] Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video
Python library to use and implement packages in OptunaHub
The registry of the OptunaHub packages
RTSS / RivaTuner Overlay
Benchmark PHP, HHVM and Zephir
A component render time benchmarking suite for React
Attention-based View Selection Networks for Light-field Disparity Estimation
:zap: A collection of common functions for Fiber with better performance, fewer allocations, and fewer dependencies.
Video dataset dedicated to portrait-mode video recognition.
Collection of tips for faster spatial data processing in R
[NAACL 2025 Oral] Multimodal Needle in a Haystack (MMNeedle): Benchmarking Long-Context Capability of Multimodal Large Language Models
GraphQL benchmarks using the-benchmarker framework.
A Benchmark Suite for Heterogeneous System Computation
🏋️ Bash Script which runs several Linux benchmarks (Sysbench, UnixBench and Geekbench)
This is a lightweight GAN developed for real-time deblurring. The model has a super tiny size and a rapid inference time. The motivation is to boost m...
[ICML 2023] This project is the official implementation of our accepted ICML 2023 paper BiBench: Benchmarking and Analyzing Network Binarization.
A comprehensive benchmark for evaluating deep research agents on academic survey tasks
[ICLR 2025] MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
A benchmarking tool for testing and comparing the performance of both embedded and networked SQL and NoSQL databases.
Text style transfer benchmark
A tiny, zero-dependency utility for measuring code execution time in high-resolution real time. Works in Node.js, browsers, Deno and Bun.
Testing different approaches to improve PHP script performance
Tests and benchmarks for cudnn (and in the future, other nvidia libraries)
This repository contains the code base for the Open Stream Processing Benchmark.
a simple benchmark testing tool implemented in golang with some small features
Benchmark your local LLMs.
Show how to perform fast retraining with LightGBM in different business cases
Realtime benchmarks for PHP code