Measurement integrity
Drift, reliability, missingness and calibration — whether the number is fit to decide on at all.
28 of 388 tools.
Analyze deep uncertainty minimax regret
Apply Savage minimax regret when scenario probabilities are not defensible, compare maximin and equal-weight choices, and use PRIM-style iterative peeling to discover compact context boxes where the robust choice remains vulnerable.
Audit aggregate metric reversal
Detect Simpson's-paradox-style sign reversals between an executive aggregate relationship and its weighted within-stratum fixed-effect relationship, with whole-stratum bootstrap uncertainty and practical-magnitude gates.
Audit informative metric missingness
Audit whether aggregate metric availability is associated with a governed outcome using permutation inference, bootstrap intervals, practical effect gates, and false-discovery control.
Audit joint metric dependency drift
Detect changes in cross-metric dependence with empirical-copula ranks, random-feature permutation inference, sliced Wasserstein magnitude, and FDR-controlled pair diagnostics.
Audit multivariate metric drift
Detect material distribution shifts with reference-fixed quantile bins, PSI, Jensen-Shannon divergence, standardized Wasserstein distance, permutation tests, and FDR control.
Audit point in time model integrity
Gate an analytical or AI model on point-in-time correctness by auditing actual feature availability, snapshot creation, target-window ordering, outcome resolution, source-record reuse, and embargoed train/calibration/test boundaries, with row and feature diagnostics rather than a generic leakage warning.
Audit policy feedback performativity
Audit whether deploying a probability-driven policy is associated with a changed score-to-outcome relationship: compute cluster-level exposed-versus-comparison pre/post differences in predictions, outcomes, calibration residuals, and Brier loss; bootstrap the assignment unit; and abstain when baseline balance or score overlap cannot support the comparison.
Audit probabilistic forecasts
Audit whether resolved probability forecasts are accurate, calibrated, discriminating, and better than a base-rate prediction.
Audit proxy metric integrity
Audit whether an incentivized proxy structurally decoupled from outcomes or harmed guardrails using counterfactual residuals, bootstrap break tests, placebos, and multiplicity correction.
Audit selective label partial identification
Partially identify event risk, calibration gap, and Brier score when a policy selectively reveals outcomes: retain missing labels, model observed-versus-missing event odds within each decision stratum under a governed sensitivity ratio, propagate Beta posterior uncertainty, expose unsupported strata and label coverage, and fail closed on wide bounds or undocumented decision rules.
Compute robust operating viability kernel
Compute the maximal robust controlled-invariant set of safe operating states under complete set-valued state-action transitions, identify every feedback action that keeps all modeled successors viable indefinitely, and expose finite guaranteed-survival layers for states outside the kernel without using probabilities or rewards.
Conformalize prediction intervals
Apply finite-sample split-conformal inflation to model intervals, with Mondrian group corrections and explicit global fallback for sparse groups.
Control online alert false discoveries
Control false discoveries across a prespecified live hypothesis stream with LORD++ and an infinite geometric alpha-spending sequence.
Discover environment invariant predictive model
Search every nonempty subset of up to eight candidate features for a sparse predictive relationship whose validation residual bias and error remain within governed limits across represented environments, select without touching the test split, and compare the chosen model once against the full model on future-held-out environment data.
Estimate multilevel metric generalizability
Decompose aggregate management-metric variance into unit, period, and residual components, bootstrap reliability, and calculate the sampling needed for dependable comparisons.
Fit anchor regression shift robust model
Fit anchor regression across declared operating environments, penalizing residual variation predictable from environment anchors over a governed gamma path; choose robustness strength only on held-out worst-environment RMSE; and expose average fit, environment bias, coefficients, and leave-one-environment stability without claiming generic or causal invariance.
Fit cross fitted isotonic recalibrator
Repair monotone probability calibration with pool-adjacent-violators while using cross-fitting and a paired bootstrap to prove out-of-sample Brier improvement.
Forecast change adoption bass diffusion
Forecast aggregate organizational change or tool adoption with a Bayesian Bass diffusion model learned from reconciled historical cohorts, jointly estimating spontaneous innovation and imitation, simulating posterior uptake under per-cohort enablement capacity, pricing enabled value, exposing grid-boundary misspecification, and gating a target adoption probability.
Forecast engineering investment benefit realization
Forecast whether an engineering-investment portfolio will realize finance-defined benefits within a decision horizon using a partially pooled Bayesian hurdle/lognormal model for zero-benefit risk, positive benefit multiples, and realization lag; correlated organization shocks; discounting; NPV/ROI gates; and explicit unseen-category fallback.
Forecast intervention effect half life
Learn how quickly a governed intervention's effect decays across resolved cohorts using a shared exponential half-life, cohort-specific amplitudes, a persistent floor, reported standard errors, and a profiled Bayesian grid; then forecast effect/value paths and when each current intervention is likely to fall below a practical threshold.
Monitor forecast calibration eprocess
Continuously monitor binary forecasts for calibration drift with an anytime-valid mixture e-process that does not incur a repeated-peeking penalty.
Optimize alert decision threshold
Choose a cost-sensitive alert action threshold using cross-validated decision curves and bootstrap net-benefit evidence against constant policies.
Optimize stratified evidence sampling
Allocate a fixed evidence budget across finite-population strata with exact discrete Neyman allocation and quantify precision gained over proportional sampling.
Optimize technical debt paydown portfolio
Choose a dependency- and exclusion-safe technical-debt portfolio under capacity and cash budgets by discounting compounding recurring drag, failure exposure, remediation effectiveness, risk reduction, and engineering opportunity cost.
Score evidence readiness
Gate an analytical claim on coverage, freshness, identity resolution, sample size, and source agreement.
Stack resolved probability forecasts
Fit convex weights to frozen probability forecasts on chronological training history and require bootstrap-validated log-loss improvement over the training-selected best component on future outcomes.
Validate temporal leading indicators
Validate aggregate leading indicators only when their lagged history improves expanding-window forecasts beyond target autoregression, with block inference and FDR.
Value of information
Calculate how much it is worth paying for more information before making an engineering decision.