--- title: "Validation Before Promotion: Recovery, SBC, PPC, and Transportability" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Validation Before Promotion: Recovery, SBC, PPC, and Transportability} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") library(eyeprocess) ``` # A common validation contract Every advanced model should answer the same questions before promotion: * can its parameters be recovered under its own generator? * how biased and noisy are the estimates? * do intervals cover at their nominal rate? * how often does estimation fail? * are parameters empirically identifiable? * is Bayesian computation calibrated? * can replicated data reproduce scientifically important features? * what happens under misspecification and preprocessing changes? * does measurement transport to held-out devices, sessions, sites, and items? * does the process channel add out-of-sample information beyond conventional response/RT information? ```{r} spec <- eyeprocess::irt_validation_spec( model_id = "joint_gaze_rt", replications = 500, parameters = c("ability", "speed", "engagement"), grouped_validation = c("device", "session", "site") ) spec ``` # Recovery table The canonical recovery table has one row per replicate/parameter and columns `replicate`, `parameter`, `truth`, `estimate`, and optionally `lower`, `upper`, `converged`, `scenario`, `engine`, and `failure_type`. ```{r, eval=FALSE} summary <- summarize_parameter_recovery(recovery) audit_bias(summary, threshold = .10) audit_rmse(summary, threshold = .30) audit_coverage(summary, minimum = .90) audit_interval_width(summary) audit_convergence(recovery, minimum = .95) audit_identifiability(recovery) validation_mcse(recovery, "coverage") recommended_validation_replications(target_mcse = .01, metric = "coverage") plot(summary) ``` # Prior SBC ```{r, eval=FALSE} sbc <- run_sbc( simulator = function(r) simulate_one_dataset(r), fitter = function(dat) fit_bayesian_model(dat), posterior_draws = function(fit) as.matrix(fit$draws), replications = 250 ) audit_sbc(sbc) plot(sbc, parameter = "ability_sd") ``` SBC is an inference-algorithm validation. It does not establish empirical model fit or construct validity. # Posterior SBC Posterior SBC addresses a different question: whether inference is calibrated in the region relevant **conditional on the observed data**. Because the correct conditional self-consistency experiment is model-specific, eyeprocess requires an explicit callback rather than pretending ordinary SBC is posterior SBC. ```{r, eval=FALSE} contract <- posterior_sbc_contract(function(replicate, observed_data) { # Model-specific implementation following the posterior-SBC construction. # Must return the simulated truth and posterior draws from the corresponding # conditional self-consistency experiment. list(truth = truth, draws = draws) }) psbc <- run_posterior_sbc(observed_data, contract, replications = 100) audit_sbc(psbc) ``` # Posterior predictive checks ```{r, eval=FALSE} posterior_predictive_discrepancies( observed = observed_fixations, replicated = replicated_fixations, discrepancies = list( mean = mean, sd = sd, zero_rate = function(x) mean(x == 0), p95 = function(x) unname(quantile(x, .95)) ) ) ``` Choose discrepancies that could falsify the scientific use of the model: tail fixation counts, omission rate, transition entropy, pupil peak timing, response accuracy by item difficulty, and other substantively meaningful summaries. # Misspecification ```{r, eval=FALSE} stress_test_latent_distribution(runner) stress_test_local_dependence(runner) stress_test_speededness(runner) stress_test_missingness(runner) stress_test_preprocessing(runner, variants = c( "default", "strict_validity", "alternate_fixation_detector", "alternate_pupil_filter" )) ``` The `runner` owns data generation and fitting; the framework records scenario, replicate, results, and classified failures. # External and grouped validation ```{r, eval=FALSE} external_validate_irt(train, external, fitter, predictor, scorer) leave_device_out_validation(data, "device", fitter, predictor, scorer) leave_session_out_validation(data, "session", fitter, predictor, scorer) leave_site_out_validation(data, "site", fitter, predictor, scorer) leave_item_out_validation(data, "item_id", fitter, predictor, scorer) ``` Then quantify transportability rather than reporting only a pooled score: ```{r, eval=FALSE} audit_measurement_transportability( held_out_results, metric = "rmse", higher_is_better = FALSE, max_range = .15 ) ``` # Incremental information and negative controls A process channel should survive a stricter test than in-sample significance. ```{r, eval=FALSE} inc <- audit_channel_incremental_information( data, fold = "participant_fold", baseline_fitter = fit_response_rt, process_fitter = fit_response_rt_gaze, predictor = predict_trait, scorer = trait_rmse, higher_is_better = FALSE ) plot(inc) negative_control_process_test( data, process_columns = c("fixation_count", "evidence_dwell"), within = c("person_id", "item_id"), evaluator = full_crossvalidated_score, permutations = 250, higher_is_better = TRUE ) ``` This guards against adding gaze merely because a high-dimensional channel can improve training fit. # Process-dependent discrimination ```{r, eval=FALSE} pd <- process_dependent_discrimination_audit( data, response = "correct", theta = "theta", process = "rt_ms", person = "person_id", item = "item_id" ) plot(pd) ``` The diagnostic asks whether effective response discrimination varies with the person-by-item process residual. It should be described as an association unless a design identifies a causal mechanism. # Evidence grade ```{r, eval=FALSE} grade_model_evidence( recovery = recovery, spec = spec, external_validation = held_out_results, sbc = sbc, ppc = ppc, semantic_roundtrip = roundtrip ) ``` The grade is a governance summary of supplied evidence, not a substitute for construct validity or independent replication.