Open, careful discovery — AI helpers that show their work.Open, measured discovery — agents that show their work.
Public day-by-day notes. Claims need citations. Uncertainty is labeled. This is a public preview — measured claims only; proved yet: no until honesty gates clear.
Public chronological notes. Claims require citations. Uncertainty is labeled. Public research preview — measured claims only; support_claim false until gates clear.
2026-09-05 · Track 1 — formula test partway (380 of 1920)Prop-01 E1-reduced partial
FLUX ENGINE checkpoint from the live formula bake-off files only. Partway through — 380 of 1920 finished. Proved yet: no. Notable progress — not an official win; not a breakthrough.
Plan for this run: 8 known laws; sample sizes 50/100/400; noise 0 or 5%; seeds 0…19; time-padding off; 1920 runs planned.
In this checkpoint: only linear-decay + harmonic-oscillator so far. Accepts: A0 200/200; A4 177/180 (3 rejects). Empty-equation fails: 0.
Matched timing (checkpoint): average run times within about 20% — yes (≈12.09 s vs ≈11.24 s). Pairwise within-band 0.733 (132/180).
Example (no noise, 50 points, 20 seeds): linear recovery 1.000 both tools; harmonic 1.000 / 0.950. Mid-noise and other laws still incomplete here.
Limits: Proxy tables untouched. Don’t treat partial rates as the full-grid success check. Wait for all 1920 before any “proved” language.
FLUX ENGINE E1-reduced live PySR/PhySO checkpoint from sims/sr-ladder/results/prop01_live_e1_full.md / .json only. partial=true — 380/1920 cells. prop01_support_claim=False (note: “Partial checkpoint — claim stays false”). Notable progress — not §5 support; not breakthrough.
Grid (this sitting): 8 laws planned; N∈{50,100,400}; σ∈{0.0,0.05}; seeds 0…19; soft-pad OFF; planned 1920 cells (drop N=200 vs original 2560).
In this checkpoint: records cover linear_decay + harmonic_oscillator only so far. Accepts: A0 200/200; A4 177/180 (3 rejects). A4 zero-eq cells: 0.
Matched wall (checkpoint): arm means within ±20% band — True (A0 mean ≈12.09 s, A4 ≈11.24 s, rel_gap ≈0.070). Paired within-band frac 0.733 (132/180).
Example aggregates (σ=0, N=50, 20 seeds): linear A0/A4 recovery 1.000; harmonic A0 1.000, A4 0.950. Mid-noise and remaining laws still incomplete in this partial file.
Limits: STLS untouched. Do not treat partial rates as full-grid §5. Wait for non-partial 1920 before any support language.
Also see prop01_live_e1_full_reduction.md. Feed: Live Findings.
2026-09-05 · Track 2 — chemistry-family ban comboProp-02 Option B chemsys combo
PRISM GRID Option B test (diversity + chemistry-family ban) from the measured write-up only. Proved yet: no. Fingerprint path is a side check — not the main graph-network proof.
No ban: uncertainty 0.2925; random 0.1072; diversity 0.0516 (diversity still worse than random).
Ban seen chemistry families: uncertainty 0.2950; random 0.0982; diversity 0.0516; leftover seed-family share → 0.
Read: Smart uncertainty picking still beats random with or without the ban; diversity stays worst.
Limits: Rough check only; exploratory Option B.
PRISM GRID Option B diversity + chemsys-ban combo from data/processed/track2-mp-bandgap/optionb_diversity_leakage.md only. prop02_success_claimed=False. Fingerprint secondary — not GNN primary support.
No ban: U 0.2925 ± 0.0136; R 0.1072 ± 0.0099; D 0.0516 ± 0.0074 (diversity still < random).
Ban seen chemsys: U 0.2950 ± 0.0105; R 0.0982 ± 0.0063; D 0.0516 ± 0.0074; seed-chemsys share → 0 for U/R/D.
Read: Uncertainty stays ahead of random with or without chemsys ban; diversity remains worst. Chemsys proxy does not kill the U lift on this setup.
Limits: Coarse leakage proxy; exploratory Option B only.
Figure: figures/track2_optionb_diversity_leakage.png · Live Findings.
Last mirrored / as of:loading… · This page is a human-curated archive and can lag. Live Findings is the live measured feed.This page is a human-curated archive and can lag. Live Findings (live/findings.json) is the live measured feed.
PRISM GRID calibrated-uncertainty probe on Magpie-lite (5,000 examples). Numbers from the probe write-up only. Proved yet: no. Notable — not a breakthrough; fingerprint side path only.
PRISM GRID calibrated-uncertainty probe on Magpie-lite N=5k. Numbers from data/processed/track2-mp-bandgap/structure_al_ucal_probe.md only. prop02_success_claimed=False. Notable — not breakthrough; fingerprint secondary only.
Calibrated uncertainty: final recall 0.3283 ± 0.0145.
Raw uncertainty: final recall 0.2914 ± 0.0133.
Random: final recall 0.1043 ± 0.0059.
Method: calibrate hit chances, then pick examples the model is unsure about. Setup: 5,000 examples, batch 25, budget 400, 5 seeds.
Limits: Simple probe — not the full materials stack; proved yet: no (main graph method still required).
Random (R): final recall 0.1043 ± 0.0059, CI95 [0.0992, 0.1094]; AULC 0.0627 ± 0.0064 (~0.104).
Method: U-raw = RF vote-std; U-cal = isotonic map of mean P→label on 30% labeled holdout, acquire by entropy of calibrated P(hit). Setup N=5000, seed=100, batch=25, budget=400, 5 seeds.
Limits: Simple isotonic probe — not materialsUQ full stack; not Prop-02 support (GNN primary still required).
Figure: figures/track2_structure_al_ucal_probe.png · metrics: structure_al_ucal_probe_metrics.json · Live Findings.
2026-09-05 · Track 2 — diversity picking armProp-02 diversity arm
PRISM GRID farthest-point diversity arm on Magpie-lite (5,000 examples). Numbers from the diversity write-up only. Proved yet: no. Notable — not a breakthrough; not the main graph-network proof.
PRISM GRID farthest-point diversity arm on the Magpie-lite N=5k setup. Numbers from data/processed/track2-mp-bandgap/structure_al_diversity_arm.md only. prop02_success_claimed=False. Notable — not breakthrough; not GNN primary support.
Uncertainty: final recall 0.3018 ± 0.0100.
Random: final recall 0.1043 ± 0.0059.
Diversity: final recall 0.0534 ± 0.0037 — worse than random.
PRISM GRID near-duplicate check on the Magpie-lite (5,000) uncertainty setup. Numbers from the diagnostic write-up only. Proved yet: no. Notable — not a breakthrough; not a win decision.
PRISM GRID near-duplicate / leakage diagnostic on the Magpie-lite N=5k uncertainty setup. Numbers from data/processed/track2-mp-bandgap/structure_al_leakage_diag.md only. prop02_success_claimed=False. Notable — not breakthrough; not a support decision.
Unfiltered uncertainty: final recall 0.3018 ± 0.0100.
Tight similarity filters: recall stays about 0.30 — does not collapse toward random.
Pool vs seed: materials are not mostly near-copies of the starter set.
Limits: Fingerprint similarity ≠ full structure-graph similarity; proved yet: no.
Unfiltered U: final recall 0.3018 ± 0.0100, CI95 [0.2930, 0.3106].
Cosine filters stable (~0.30): cos≥0.99 → 0.3025 ± 0.0114; cos≥0.95 → 0.3039 ± 0.0089; cos≥0.90 → 0.2961 ± 0.0098. Recall does not collapse toward random under these filters.
Pool vs seed: mean max-cosine to seed 0.6278 ± 0.0048; p95 0.8852; frac cos≥0.99 / 0.95 / 0.90 among unlabeled ≈ 0.0007 / 0.0074 / 0.0370.
Limits: Fingerprint cosine ≠ structure-graph similarity; fingerprint-only path still ≠ Prop-02 support (GNN primary still required).
Figure: figures/track2_structure_al_leakage.png · metrics: structure_al_leakage_diag_metrics.json · Live Findings.
2026-09-05 · Track 1 — small live formula pilotProp-01 E1 live pilot
FLUX ENGINE small live formula pilot finished. Numbers from the pilot write-up only. Proved yet: no. Notable hygiene check — not an official win; not a breakthrough.
FLUX ENGINE E1 adapter/metric pilot finished. Numbers from sims/sr-ladder/results/prop01_live_e1_pilot.md only. prop01_support_claim=False. Notable hygiene check — not a §5 support claim; not breakthrough.
Setup: 2 laws × 3 seeds × sizes 50/100 × no noise × two tools = 24 runs. About 97.6 s total.
Recovery proxy: held-out residual < R* (not structural τ). A0 PySR: recovery 1.0 on all four law×N groups (12/12 accept). A4 PhySO: recovery 1.0 on harmonic N=50 and linear N=100; 0.667 on harmonic N=100 and linear N=50 (10/12 accept).
Matched wall-clock: target ±20% between arm means — miss. A0 mean wall_s 1.976 (n=12) vs A4 mean 6.147 (n=12). Within band: False.
Limits: STLS tables not overwritten. Full ≥20-seed E1 still required for any support flip; matched-clock fix needed first.
Artifacts: prop01_live_e1_pilot.md / .json. Feed: Live Findings.
2026-09-05 · Track 2 — which features help mostProp-02 structure ablation
Metal flag + geometry drive most of the Magpie-lite structure boost (leave-one-group-out on the 5,000-example pilot, uncertainty arm). Numbers from the ablation write-up only. Proved yet: no. Notable — not a breakthrough.
is_metal + geometry drive most of the Magpie-lite +structure lift (leave-one-group-out on the N=5k pilot, uncertainty arm). Numbers from data/processed/track2-mp-bandgap/structure_al_feature_ablation.md only. prop02_success_claimed=False. Notable — not breakthrough.
Full composition+structure: uncertainty recall 0.3018 ± 0.0100.
Drop metal/electronic: 0.2781 → down about 0.024 (largest hit).
Drop geometry: 0.2832 → down about 0.019.
Composition-only: 0.2419 (down about 0.060 vs full).
Limits: Leave-one-group-out ≠ proof of cause; proved yet: no.
Full composition+structure: U final recall 0.3018 ± 0.0100, CI95 [0.2930, 0.3106] (91 features).
Drop electronic / is_metal: 0.2781 ± 0.0172 → Δ −0.0237 vs full (largest structure-group hit).
Drop geometry (log1p_volume, density): 0.2832 ± 0.0169 → Δ −0.0186.
Drop size / thermo: Δ −0.0079 / −0.0072. Composition-only baseline: 0.2419 ± 0.0164 (Δ −0.0599 vs full).
RF impurity (1 seed, supportive): geometry group 0.1332; electronic/is_metal 0.0846; thermo group 0.0000.
Limits: LOGO ≠ causal; not a Prop-02 support decision.
Figure: figures/track2_structure_al_ablation.png · metrics: structure_al_feature_ablation_metrics.json · feed: Live Findings.