Skip to content
#

negative-results

Here are 163 public repositories matching this topic...

llama.cpp speculative decoding measured on one RTX 3090, Qwen3.6-35B-A3B UD-Q4_K_XL, commit 3737e4137. Published figures are re-derived from the committed data by a checker that fails on drift, a coverage probe reports how many of them it actually covers, and ERRATA.md lists this study's own retracted claims.

  • Updated Sep 3, 2026
  • Python

A negative result on joint-embedding predictive architectures for prediction markets. The martingale-collapse diagnostic is sound and predicts nothing, more training makes the representation worse, and the metric everything was ranked by was mostly measuring input reconstruction. Pre-registered gates, full findings record.

  • Updated Aug 13, 2026
  • Python

JEPA agent playing Minecraft from pixels: latent world model + MPC planning, 664K params on one 8GB GPU, trained on raw gameplay with no labels. A complete lab notebook - including a 20-attempt research dead end, documented with its root cause.

  • Updated Aug 10, 2026
  • HTML

Resample or reroute after a weak-verifier stop? Pre-registered measurements of recoverable stopping debt on MBPP+, a two-sided action-support gate on BigCodeBench that fails closed, a LiveCodeBench observability ladder, and the exchangeable-actions reference showing realized-maximum gaps carry no selector signal. Artifacts for arXiv:2607.08665v3.

  • Updated Sep 4, 2026
  • Python

A global, open-source registry and standardized schema for null findings, non-significant outcomes, and failed trials in medical research. Built to eliminate publication bias and accelerate biomedical discovery.

  • Updated Jul 26, 2026
  • Python

A $0 falsification lab across two markets — crypto (~111 hypotheses) and Polymarket prediction markets (172,830 resolved markets, 1.36M trades). 184+ techniques through one committed anti-overfitting gauntlet. 0 survive to a deployable edge; the reusable validation harness is the asset. Agent-ready. MIT.

  • Updated Jun 16, 2026
  • TypeScript

Component-level benchmark for catastrophic forgetting in world models. Two negative results: forgetting does not follow the labelled task-distance axis, and it happens in the encoder, where the usual metrics cannot see it. 375 runs, with code, data and paper.

  • Updated Aug 14, 2026
  • Python

Add this topic to your repo

To associate your repository with the negative-results topic, visit your repo's landing page and select "manage topics."

Learn more