Topic

benchmark

Repositories (1866)

nagi-bench
nagi-bench nagi-studio HTML

One-shot LLM eval cases by NAGI STUDIO - same prompt, different agents (model + harness), runnable artifacts side by side.

60
chi-bench
chi-bench actava-ai Python

Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

60
argus-validation-benchmarks
argus-validation-benchmarks pensar-x Python

Self-contained, Dockerized offensive security challenges for evaluating AI-powered penetration testing agents. Covers modern tech stacks (Node.js, Pyt...

60
functional-components-benchmark
functional-components-benchmark missive JavaScript

Directly calling functional components instead of mounting them is faster.

59
benchmarkstt
benchmarkstt ebu Python

Open Source AI Benchmarking toolkit for benchmarking speech to text services

59
touchstone
touchstone lorenzwalthert R

Smart benchmarking of pull requests with statistical confidence

59
modd
modd bborja MATLAB

Dataset and Evaluation Scripts for Obstacle Detection via Semantic Segmentation in a Marine Environment

59
gosbench
gosbench mulbc Go

Distributed S3 benchmarking tool - Replacement of Cosbench

59
ucsb
ucsb unum-cloud C++

Wide NoSQL benchmark for RocksDB, LevelDB, Redis, WiredTiger and MongoDB extending the Yahoo Cloud Serving Benchmark

59
SALOD
SALOD moothes Python

A benchmark for Salient Object Detection (SOD).

59
f5-bench
f5-bench ikxin TypeScript

利用 Fetch API 向目标网站发送频繁请求,模拟按下 F5 刷新的效果,以测试服务器的资源限制。

59
BenchmarkDotNetVisualizer
BenchmarkDotNetVisualizer mjebrahimi HTML

🌈 Visualizes your BenchmarkDotNet benchmarks to Colorful images and Feature-rich HTML (and maybe powerful charts in the future!)

59
Diabetica
Diabetica waltonfuture Python

[SCI-FM@ICLR 2025] Specialized LLMs capable of handling various diabetes tasks

59
ALERT
ALERT Babelscape Python

Official repository for the paper "ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming"

59
autoresearch-automl
autoresearch-automl ferreirafabio Python

Can LLMs beat classical HPO? A benchmark comparing classical, LLM-based, and hybrid methods on Karpathy's autoresearch.

59
sycophancy
sycophancy lechmazur

LLM benchmark and leaderboard for narrator-bias sycophancy, opposite-narrator contradictions, and judgment consistency.

59
llama-optimus
llama-optimus BrunoArsioli Python

Lightweight Python tool using Optuna for tuning llama.cpp flags: towards optimal tok/s for your machine

59
kotlin_tutorial
kotlin_tutorial fengzhizi715 Kotlin

掘金的小册《Android 进阶:基于 Kotlin 的 Android App 开发实践》中的相关的例子

58
synthmark
synthmark google C++

Audio performance benchmark for jitter, theoretical latency, etc.

58
SODBenchmark
SODBenchmark DengPingFan

Salient objects in clutter, TPAMI, 2022

58
SLAM-under-Perturbation
SLAM-under-Perturbation Xiaohao-Xu C++

[ICLR 2025] Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

58
optunahub
optunahub optuna Python

Python library to use and implement packages in OptunaHub

58
optunahub-registry
optunahub-registry optuna Jupyter Notebook

The registry of the OptunaHub packages

58
Benchmark-Overlays-2
Benchmark-Overlays-2 TroyMetrics

RTSS / RivaTuner Overlay

58
Benchmark-PHP-HHVM-Zephir
Benchmark-PHP-HHVM-Zephir treffynnon Shell

Benchmark PHP, HHVM and Zephir

57
fuego
fuego apiv JavaScript

A component render time benchmarking suite for React

57
LFattNet
LFattNet LIAGM Python

Attention-based View Selection Networks for Light-field Disparity Estimation

57
utils
utils gofiber Go

:zap: A collection of common functions for Fiber with better performance, fewer allocations, and fewer dependencies.

57
Portrait-Mode-Video
Portrait-Mode-Video bytedance Python

Video dataset dedicated to portrait-mode video recognition.

57
geotips
geotips kadyb HTML

Collection of tips for faster spatial data processing in R

57
multimodal-needle-in-a-haystack
multimodal-needle-in-a-haystack Wang-ML-Lab Python

[NAACL 2025 Oral] Multimodal Needle in a Haystack (MMNeedle): Benchmarking Long-Context Capability of Multimodal Large Language Models

57
graphql-benchmarks
graphql-benchmarks the-benchmarker Ruby

GraphQL benchmarks using the-benchmarker framework.

56
Hetero-Mark
Hetero-Mark NUCAR-DEV Jupyter Notebook

A Benchmark Suite for Heterogeneous System Computation

56
benchmark
benchmark Cyclenerd Shell

🏋️ Bash Script which runs several Linux benchmarks (Sysbench, UnixBench and Geekbench)

56
Ghost-DeblurGAN
Ghost-DeblurGAN York-SDCNLab Python

This is a lightweight GAN developed for real-time deblurring. The model has a super tiny size and a rapid inference time. The motivation is to boost m...

56
BiBench
BiBench AI-Efficiency Python

[ICML 2023] This project is the official implementation of our accepted ICML 2023 paper BiBench: Benchmarking and Analyzing Network Binarization.

56
ReportBench
ReportBench ByteDance-BandAI Python

A comprehensive benchmark for evaluating deep research agents on academic survey tasks

56
MMFakeBench
MMFakeBench liuxuannan Python

[ICLR 2025] MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs

56
crud-bench
crud-bench surrealdb Rust

A benchmarking tool for testing and comparing the performance of both embedded and networked SQL and NoSQL databases.

56
text-style-transfer-benchmark
text-style-transfer-benchmark ykshi

Text style transfer benchmark

55
perfy
perfy onury TypeScript

A tiny, zero-dependency utility for measuring code execution time in high-resolution real time. Works in Node.js, browsers, Deno and Bun.

55
faster-php
faster-php devmount PHP

Testing different approaches to improve PHP script performance

55
nvidia_libs_test
nvidia_libs_test google C++

Tests and benchmarks for cudnn (and in the future, other nvidia libraries)

55
BenchmarkCI.jl
BenchmarkCI.jl tkf Julia
55
open-stream-processing-benchmark
open-stream-processing-benchmark Klarrio Jupyter Notebook

This repository contains the code base for the Open Stream Processing Benchmark.

55
benchmark
benchmark cnlh Go

a simple benchmark testing tool implemented in golang with some small features

55
benchllama
benchllama srikanth235 Python

Benchmark your local LLMs.

55
fast_retraining
fast_retraining Azure Jupyter Notebook

Show how to perform fast retraining with LightGBM in different business cases

54
PHPench
PHPench mre PHP

Realtime benchmarks for PHP code

54
m1-cpu-benchmarks
m1-cpu-benchmarks tlkh Jupyter Notebook
54