News
Latest
Top
Search
Submit
Login
Search
▲
124
Benchmarking leading AI agents against Google reCAPTCHA v2
(research.roundtable.ai)
by mdahardy |
view
|
97 comments
▲
53
Drawing Text Isn't Simple: Benchmarking Console vs. Graphical Rendering
(cv.co.hu)
by PaulHoule |
view
|
41 comments
▲
31
How Good Are Chinese CPUs? Benchmarking the Loongson 3A6000
(lemire.me)
by ashvardanian |
view
|
1 comments
▲
28
Benchmarking the Most Reliable Document Parsing API
(tensorlake.ai)
by calavera |
view
|
14 comments
▲
10
Benchmarking NVENC video transcoding on the Pi
(jeffgeerling.com)
by ingve |
view
|
0 comments
▲
5
Benchmarking KDB-X vs. QuestDB, ClickHouse, TimescaleDB and InfluxDB
(kx.com)
by rustc |
view
|
0 comments
▲
4
Benchmarking my Redis clone in Zig (a web dev learning systems)
(charlesfonseca.substack.com)
by barddoo |
view
|
1 comments
▲
3
CodSpeed CLI: Deterministic benchmarking for any executable
(github.com)
by art049 |
view
|
0 comments
▲
3
Benchmarking GPT-5.1 vs. Gemini 3.0 vs. Opus 4.5 across 3 Coding Tasks
(blog.kilo.ai)
by heymax054 |
view
|
0 comments
▲
3
Benchmarking LLMs at the Frontier of Physics
(artificialanalysis.ai)
by mustaphah |
view
|
0 comments
▲
3
Benchmarking Language Implementations: Am I doing it right? Get Early Feedback
(stefan-marr.de)
by speckx |
view
|
0 comments
▲
3
Powering AI at Scale: Benchmarking 1B Vectors in YugabyteDB
(yugabyte.com)
by ashvardanian |
view
|
0 comments
▲
3
Benchmarking the Cost of Java's EnumSet – A Second Look
(kinnen.de)
by birdculture |
view
|
0 comments
▲
3
Benchmarking multilingual long-context language models
(arxiv.org)
by sysoleg |
view
|
0 comments
▲
2
Benchmarking Speculative Decoding on DGX Spark: 11.5 → 29.5 tok/s
(alephinitesimal.com)
by Alephinitesimal |
view
|
0 comments
▲
2
VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection
(arxiv.org)
by wslh |
view
|
0 comments
▲
2
Show HN: PantheonGPU – GPU health testing and AI workload benchmarking
(pantheongpu.com)
by saqibkhan1992 |
view
|
0 comments
▲
2
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
(arxiv.org)
by yruzin |
view
|
0 comments
▲
2
ClickCannon: Building a Tool for Benchmarking ClickHouse
(clickhouse.com)
by mikeshi42 |
view
|
0 comments
▲
2
FlowerBench: Benchmarking AI Agents on Real Enterprise Work
(flower.ai)
by dimitrisflwr |
view
|
1 comments
▲
2
Giving a domain a hill to climb: benchmarking as data activation
(sparsethought.com)
by galsapir |
view
|
0 comments
▲
2
Ask HN: Is there a recognized standard for swarm intelligence benchmarking?
by stephanieriggs |
view
|
0 comments
▲
2
Reverse Benchmarking
(dominiknitsch.com)
by wseqyrku |
view
|
0 comments
▲
2
Benchmarking node collision algorithms for React/Svelte Flow
(xyflow.com)
by moklick |
view
|
0 comments
▲
2
Show HN: Benchmark-ips-Python – benchmarking tool for Python
(github.com)
by Igor_Wiwi |
view
|
0 comments
▲
2
Benchmarking Checksum Tools
(heitorpb.github.io)
by furkansahin |
view
|
0 comments
▲
2
Dell Pro Max with GB10 Arrives for Linux Performance Benchmarking Review
(phoronix.com)
by rbanffy |
view
|
0 comments
▲
2
Benchmarking the Thomson Reuters legal agent
(thomsonreuters.com)
by gk1 |
view
|
0 comments
▲
2
Benchmarking the AMD EPYC 9V64H: Azure HBv5's Custom AMD CPU with HBM3
(phoronix.com)
by ashvardanian |
view
|
0 comments
▲
1
Benchmarking LLMs' Swarm Intelligence
(github.com)
by novia |
view
|
0 comments
▲
1
Public WebGPU Benchmarking powered by vGPU
(web3dsurvey.com)
by bhouston |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Nanobench: A simple and fast single-header microbenchmarking library for C++
(nanobench.ankerl.com)
by Almondsetat |
view
|
0 comments
▲
1
Practice Makes (Im)Perfect: Benchmarking Microarchitectural Side-Channel Attacks
(arxiv.org)
by torsi0n |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Fleet: GPU Benchmarking in the Browser
(webgpu-kernels-fleet.hf.space)
by theanonymousone |
view
|
0 comments
▲
1
Build a GitHub Runners Benchmarking Website
(runnerbench.com)
by dev-work-aint-e |
view
|
1 comments
▲
1
Benchmarking Confidential Computing Performance on Nvidia Blackwell GPUs
(arxiv.org)
by Jimmc414 |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Benchmarking – Frontier models go out of their way to cheat
(tuneloop.io)
by behat |
view
|
0 comments
▲
1
Benchmarking Vector Indexes
(percona.com)
by lbw1215 |
view
|
0 comments
▲
1
Benchmarking Pocket-Scale Inference
(artificialanalysis.ai)
by sys42590 |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Pipette: A benchmarking suite for on-device intelligence
(liquid.ai)
by Philpax |
view
|
0 comments
▲
1
Intelligence at pocket scale: Benchmarking small models and mobile phones
(artificialanalysis.ai)
by pember |
view
|
0 comments
▲
1
Benchmarking Agentgateway vs. LiteLLM's Rust Mode
(agentgateway.dev)
by j0selit0 |
view
|
0 comments
▲
1
Vast.ai Benchmarking – Detect/Log Garbage Servers
(github.com)
by icelancer |
view
|
0 comments
▲
1
Pitfalls of Benchmarking on Modern Systems
(stefan-marr.de)
by matt_d |
view
|
0 comments
▲
1
Benchmarking Small Models for DigiKam's Natural Language Search
(srirupa19.github.io)
by oever |
view
|
0 comments
▲
1
Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
(arxiv.org)
by tcp_handshaker |
view
|
0 comments