Pinned Loading
-
demand-forecasting-backtest
demand-forecasting-backtest PublicHonest forecast evaluation: rolling-origin backtesting, baselines that must be beaten, and quantile forecasts scored on interval calibration, not just point accuracy
Python
-
event-metrics-reconciler
event-metrics-reconciler PublicReconciling the same business metrics from two lossy event sources: canonical definitions, dirty-data guards, MAX-merge without double-counting, and invariant checks
Python
-
sherloq-mock
sherloq-mock PublicAgentic SQL co-pilot for BI dashboards: checkpoint-driven human-in-the-loop workflow, retrieval-grounded query generation, read-only guardrails, row-level ACL, budget-guarded LLM calls
Python
-
g2p-benchmark-harness
g2p-benchmark-harness PublicGrapheme-to-phoneme evaluation harness on CMUdict: WER/PER metrics, model-agreement confidence scoring, and a human-in-the-loop verification queue where targeted labeling beats random
Python
-
probabilistic-forecast-scoring
probabilistic-forecast-scoring PublicDistributional forecasts scored with CRPS and PIT calibration, plus block-bootstrap tests that separate leaderboard skill from luck
Python
-
name-pronunciation-ranker
name-pronunciation-ranker PublicContext-aware ranking of name pronunciations: locale priors, Bayesian updating from confirmations, person-level pinning, and confidence-gated serving. Hand-curated ARPAbet/IPA seed lexicon
Python
If the problem persists, check the GitHub status page or contact support.
