You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
llama.cpp speculative decoding measured on one RTX 3090, Qwen3.6-35B-A3B UD-Q4_K_XL, commit 3737e4137. Published figures are re-derived from the committed data by a checker that fails on drift, a coverage probe reports how many of them it actually covers, and ERRATA.md lists this study's own retracted claims.
Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
A negative result on joint-embedding predictive architectures for prediction markets. The martingale-collapse diagnostic is sound and predicts nothing, more training makes the representation worse, and the metric everything was ranked by was mostly measuring input reconstruction. Pre-registered gates, full findings record.
JEPA agent playing Minecraft from pixels: latent world model + MPC planning, 664K params on one 8GB GPU, trained on raw gameplay with no labels. A complete lab notebook - including a 20-attempt research dead end, documented with its root cause.
One fixed 9-stage QSAR and conformal-prediction drug-discovery pipeline run unchanged across ENPP1, NLRP3, TYK2 and IRAK4 — its most valuable outputs are the two refusals.
A reproducible benchmark: do AI coding agents violate architectural import-boundary rules inferred from real codebases? Harness, pre-registration, data, and a null result.
Resample or reroute after a weak-verifier stop? Pre-registered measurements of recoverable stopping debt on MBPP+, a two-sided action-support gate on BigCodeBench that fails closed, a LiveCodeBench observability ladder, and the exchangeable-actions reference showing realized-maximum gaps carry no selector signal. Artifacts for arXiv:2607.08665v3.
A global, open-source registry and standardized schema for null findings, non-significant outcomes, and failed trials in medical research. Built to eliminate publication bias and accelerate biomedical discovery.
Does cache-aware MoE routing degrade generation, and would the usual metrics notice? 432 pre-registered generations. Every number traced to a result file.
A $0 falsification lab across two markets — crypto (~111 hypotheses) and Polymarket prediction markets (172,830 resolved markets, 1.36M trades). 184+ techniques through one committed anti-overfitting gauntlet. 0 survive to a deployable edge; the reusable validation harness is the asset. Agent-ready. MIT.
Research platform for discovering and rigorously falsifying crypto trading strategies: event-sourced paper trading, implementation-parity verification, and a documented negative-results record.
Component-level benchmark for catastrophic forgetting in world models. Two negative results: forgetting does not follow the labelled task-distance axis, and it happens in the encoder, where the usual metrics cannot see it. 375 runs, with code, data and paper.
Does gnomAD LOEUF constraint predict drug-target safety? A negative result: LOEUF measures genetic loss-of-function tolerance well and clinical safety of inhibition poorly.