--- title: "Manuscript-Ready Reporting Examples" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Manuscript-Ready Reporting Examples} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") library(contentvalidR) read_example <- function(name) { utils::read.csv( system.file("extdata", name, package = "contentvalidR"), stringsAsFactors = FALSE ) } ``` ## Purpose This vignette provides reporting scaffolds for the three flagship workflows. The examples are intentionally conservative: statistical screening is described as **evidence for retention or review**, not as proof that an item or scale is content valid. Final decisions should also consider construct-domain coverage, item wording, qualitative judge feedback, and the intended use of the measure. The bundled CSV files are synthetic, deterministic, and regenerated from `data-raw/build-example-data.R`. They are useful for reproducing the examples and for seeing the expected input shape before analyzing a new study. ## Item-sort study ```{r sort-fit} sort_dat <- read_example("sort_example.csv") sort_fit <- sort_validity(sort_dat) sort_sum <- summary(sort_fit) sort_fit$results sort_fit$scale_summary ``` ### Methods scaffold Report who completed the sort, how the construct definitions were presented, the available assignment alternatives, and the a priori screening rule. A concise methods statement can follow this structure: > Candidate items were evaluated in an item-sort pretest in which judges > assigned each item to the construct definition that best represented its > content. We quantified definitional correspondence using the proportion of > substantive agreement (Psa) and definitional distinctiveness using the > substantive-validity coefficient (Csv). Item-level screening used the exact > target-assignment test described by Howard and Melloy (2016), with the null > target-assignment probability and alpha specified a priori. Target-scale mean > Psa and Csv were interpreted against Colquitt et al. (2019) norms only when the > judge population matched the intended use of those norms. ### Results scaffold For this bundled example, `r sort_fit$design$n_raters` judges evaluated `r sort_fit$design$n_items` items. `r sort_sum$n_supported` items met the exact screening criterion, `r sort_sum$n_review` were flagged for review, and `r sort_sum$n_insufficient` had insufficient usable assignments. A manuscript table can usually be built directly from: ```{r sort-table} sort_fit$results[c( "item", "target", "n", "n_target", "competitor", "psa", "csv", "p_value", "status", "recommendation" )] ``` Do not report `Review` as synonymous with deletion. A review flag identifies an item for substantive inspection; retaining an item for domain coverage can be a reasonable decision when that rationale is documented. ## Construct-rating study ```{r rating-fit} rating_dat <- read_example("rating_example.csv") rating_fit <- rating_validity(rating_dat, scale_min = 1, scale_max = 5) rating_sum <- summary(rating_fit) rating_fit$results rating_fit$scale_summary ``` ### Methods scaffold > Judges rated every candidate item against each focal and orbiting construct > definition using the same response scale. We summarized correspondence with > HTC and distinctiveness with HTD. Because the same judges rated the competing > definitions, item-level inference used a repeated-measures design. Planned > paired contrasts compared each item's intended definition with every orbiting > definition; omnibus Greenhouse-Geisser-corrected inference was used when > applicable. Scale-level HTC/HTD norms from Colquitt et al. (2019) were treated > as empirical benchmarks rather than universal item cutoffs. ### Results scaffold The example contains `r rating_fit$design$n_items` items rated by `r rating_fit$design$n_raters` judges against `r rating_fit$design$n_constructs_observed` construct definitions. `r rating_sum$n_supported` items were supported by the complete screening rule and `r rating_sum$n_review` were flagged for review. ```{r rating-table} rating_fit$results[c( "item", "target", "n_complete", "strongest_competitor", "htc", "htd", "p_value", "max_contrast_p", "status", "recommendation" )] ``` For review items, report the strongest orbiting competitor. That information turns a generic statement about weak distinctiveness into a specific diagnostic about where construct overlap may be occurring. ## Expert-panel study ### Relevance ```{r expert-relevance} expert_rel <- read_example("expert_relevance_example.csv") expert_rel_matrix <- as.matrix(expert_rel[setdiff(names(expert_rel), "expert")]) expert_fit <- expert_validity( expert_rel_matrix, mode = "relevance", lo = 1, hi = 4 ) expert_sum <- summary(expert_fit) expert_fit$results expert_fit$scale_summary ``` > Experts rated the relevance of each candidate item on a bounded ordinal > scale. We summarized relevance using Aiken's V with Penfield-Giacobbi score > confidence intervals and calculated I-CVI with Polit-Beck-Owen modified kappa. > S-CVI/Ave and S-CVI/UA were reported at the scale level. Panel-size CVI > guidelines were used as review aids and were considered alongside written > expert feedback and construct coverage. In this example, `r expert_fit$design$n_judges` experts evaluated `r expert_fit$design$n_items` items. The workflow identifies `r expert_sum$n_supported` supported items and `r expert_sum$n_review` review items under its quantitative rules. ### Essentiality ```{r expert-essentiality} expert_ess <- read_example("expert_essentiality_example.csv") expert_ess_matrix <- as.matrix(expert_ess[setdiff(names(expert_ess), "expert")]) ess_fit <- expert_validity(expert_ess_matrix, mode = "essentiality") ess_fit$results ``` For an essentiality task, state that CVR is tied to a different expert judgment than relevance. Report the effective panel size and exact critical essential count for each item rather than borrowing a relevance/CVI threshold. ### Congruence ```{r expert-congruence} expert_ioc <- read_example("expert_congruence_example.csv") ioc_fit <- expert_validity(expert_ioc, mode = "congruence") ioc_fit$results ``` For IOC, report the intended objective, target IOC, strongest competing objective, and target-minus-competitor margin. The margin is diagnostic evidence about alignment; it is not a newly invented significance test. ## Minimum reproducibility statement At minimum, a manuscript or supplement should identify the package version and record the analysis settings that determine the results. For a fitted workflow: ```{r reproducibility} packageVersion("contentvalidR") sort_fit$settings sort_fit$design ``` A strong reproducibility supplement should also archive the item wording, construct definitions, judge instructions, anonymized response data when permitted, and the script that reproduces all tables and figures. ## Language to avoid Avoid claims such as "the scale was proven content valid" or "items failing the cutoff were invalid." The workflows quantify evidence from a defined pretest. They do not replace a theory-based definition of the construct domain or the researcher's responsibility to justify substantive item decisions.