--- title: "Interview-Based Catch Estimation" output: rmarkdown::html_vignette: highlight: null vignette: > %\VignetteIndexEntry{Interview-Based Catch Estimation} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) ``` ## Introduction This vignette extends the effort estimation workflow covered in "Getting Started with tidycreel" to interview-based catch and harvest estimation. The complete workflow involves five steps: 1. **Design**: Define your survey calendar and stratification 2. **Counts**: Attach instantaneous count observations 3. **Interviews**: Attach complete trip interview data 4. **CPUE**: Estimate catch per unit effort 5. **Total Catch**: Combine effort and CPUE estimates The package uses ratio-of-means estimation for catch per unit effort (CPUE), which is appropriate for access point surveys with complete trip interviews. This estimator accounts for the correlation between catch and effort within each interview. ## Survey Design and Count Data We start with the same design and count data workflow from the "Getting Started" vignette: ```{r design_counts} library(tidycreel) # Load example calendar and counts data(example_calendar) data(example_counts) # Create design with counts design <- creel_design(example_calendar, date = date, strata = day_type) design <- add_counts(design, example_counts) # Estimate total effort effort_est <- estimate_effort(design) print(effort_est) ``` The effort estimate provides the total angler-hours for the survey period, which we'll combine with CPUE to estimate total catch. ## Adding Interview Data Next, we attach interview data containing catch and effort for complete trips: ```{r interviews} # Load example interview data data(example_interviews) head(example_interviews) # Attach interviews to the design design <- add_interviews(design, example_interviews, catch = catch_total, effort = hours_fished, n_anglers = n_anglers, harvest = catch_kept, trip_status = trip_status, trip_duration = trip_duration ) print(design) ``` The `add_interviews()` function maps the interview data columns to the design structure. The design now shows both count and interview data attached. Interview data is treated as a parallel data stream to count data—the two datasets do not need to align on specific dates. ## Estimating CPUE With interview data attached, we can estimate catch per unit effort: ```{r cpue} # Estimate CPUE cpue_est <- estimate_catch_rate(design) print(cpue_est) ``` The CPUE estimate uses a ratio-of-means estimator: total catch divided by total effort across all interviews. This estimator is appropriate for access point surveys because it accounts for the correlation between catch and effort within each trip. The variance is computed using the delta method, accounting for the covariance between numerator and denominator. For details on the ratio-of-means formula and variance calculation, see `?estimate_catch_rate`. ## Estimating Total Catch We can combine the effort and CPUE estimates to compute total catch: ```{r total_catch} # Estimate total catch total_catch_est <- estimate_total_catch(design) print(total_catch_est) ``` The total catch estimate multiplies the effort and CPUE estimates. The variance is computed using the delta method with the formula: Var(E × C) = E² Var(C) + C² Var(E) This formula assumes independence between the count and interview data streams, which is appropriate since they are collected through separate sampling processes. See `?estimate_total_catch` for more details on the variance propagation. ## Estimating Harvest The package distinguishes between total catch (all fish caught) and harvest (fish kept). We can estimate harvest per unit effort (HPUE) and total harvest: ```{r harvest} # Estimate HPUE hpue_est <- estimate_harvest_rate(design) print(hpue_est) # Estimate total harvest total_harvest_est <- estimate_total_harvest(design) print(total_harvest_est) ``` Harvest estimation uses the same ratio-of-means approach as CPUE, but with the harvest column (`catch_kept`) as the numerator. As expected, total harvest is lower than total catch since not all caught fish are kept. ## Grouped Estimation Like effort estimation, CPUE and total catch functions support grouped estimation using the `by` parameter: ```{r grouped, eval=FALSE} # Estimate CPUE by day_type (not evaluated - example data too small) cpue_by_day <- estimate_catch_rate(design, by = day_type) # Estimate total catch by day_type total_catch_by_day <- estimate_total_catch(design, by = day_type) ``` Note: The example data has insufficient interview sample sizes for grouped estimation (weekend n=9, weekday n=13). Ratio estimation requires at least 10 observations per group for stability. In real surveys, aim for at least 30 interviews per stratum for reliable grouped estimates. ## Variance Methods Like effort estimation, CPUE estimation supports multiple variance methods. The default is Taylor linearization, but bootstrap and jackknife are also available: ```{r variance} # Bootstrap variance estimation set.seed(42) # For reproducibility cpue_boot <- estimate_catch_rate(design, variance = "bootstrap") # Jackknife variance estimation cpue_jk <- estimate_catch_rate(design, variance = "jackknife") ``` All three methods produce similar results for these data. Use bootstrap or jackknife when you want to verify Taylor linearization assumptions or when working with complex grouped estimates. **When to use each method:** - **Taylor linearization** (default): Computationally efficient and appropriate for most smooth statistics. This is the recommended default. - **Bootstrap**: Use when working with non-smooth statistics or when you want to verify Taylor linearization assumptions. More computationally intensive. - **Jackknife**: Alternative resampling method that is deterministic (unlike bootstrap). Useful for verification or when bootstrap is too slow. ## Age Structure Estimation Age data collected during interviews can be used to estimate a pressure-weighted age distribution and design-weighted mean age for the catch or harvest. Attach age records with `add_ages()`, then call `est_age_distribution()` followed by `est_mean_age()`. ```{r ages} data(example_ages) data(example_catch) head(example_ages) # Ages are read from a subsample of the catch, so the age-class totals are # scaled onto the reported catch. Grouping by species needs this table, which # is the only record of catch per species. design <- add_catch( design, example_catch, catch_uid = interview_id, interview_uid = interview_id, species = species, count = count, catch_type = catch_type ) # Attach age data — one row per aged fish design <- add_ages( design, example_ages, age_uid = interview_id, interview_uid = interview_id, species = species, age = age, age_type = age_type ) # Estimate weighted age frequency by species ad <- est_age_distribution(design, by = species) print(ad) ``` Ages come from a subsample of the catch, so `est_age_distribution()` reports a two-phase estimate: the age composition among aged fish, scaled onto the design-estimated reported catch. The call warns when it rescales, naming the factor. Shares (`percent`) are unaffected by how many fish were aged. Each row is one integer age class. The `estimate` column is the survey-design-weighted count of fish at that age, `percent` is the within-species proportion, and `cumulative_percent` accumulates from the youngest age class up. The same `variance`, `type`, and `conf_level` arguments available in `est_length_distribution()` are accepted here. ```{r mean_age} # Compute weighted mean age from the distribution object est_mean_age(ad) ``` `est_mean_age()` applies the ratio estimator (Σ a N̂_a / N̂) to the weighted age counts returned by `est_age_distribution()` and propagates variance via the delta method. ## Complete Workflow Example The full pipeline in one workflow is: ```{r complete} # Load data data(example_calendar) data(example_counts) data(example_interviews) # Build design and attach data complete_design <- creel_design(example_calendar, date = date, strata = day_type) |> add_counts(example_counts) |> add_interviews(example_interviews, catch = catch_total, effort = hours_fished, n_anglers = n_anglers, harvest = catch_kept, trip_status = trip_status, trip_duration = trip_duration ) # Compute all estimates effort <- estimate_effort(complete_design) cpue <- estimate_catch_rate(complete_design) hpue <- estimate_harvest_rate(complete_design) total_catch <- estimate_total_catch(complete_design) total_harvest <- estimate_total_harvest(complete_design) # Print key results print(effort) print(total_catch) print(total_harvest) ``` This demonstrates the complete v0.2.0 workflow from survey design through total catch and harvest estimation. ## Working with Incomplete Trips This vignette demonstrates the default workflow using complete trips, which is the recommended approach following Pollock et al. (1994) roving-access design principles. For situations with incomplete-trip interviews where you want to examine using them for estimation, see the **Incomplete Trip Estimation** vignette (`vignette("incomplete-trips", package = "tidycreel")`). That vignette covers: - When incomplete trip estimation is scientifically valid - How to validate incomplete trip estimates using TOST equivalence testing - Step-by-step workflow with `validate_incomplete_trips()` - Examples of passing and failing validation scenarios - Why you should NEVER pool complete and incomplete trips ::: {.alert .alert-warning} **Important:** The package defaults to complete trips only. Incomplete trip estimation requires explicit opt-in via the `use_trips` parameter in `estimate_catch_rate()` and should only be used after validation with `validate_incomplete_trips()`. ::: ## Next Steps For more details on interview-based estimation functions, see: - `?add_interviews` - Attach interview data to a design - `?estimate_catch_rate` - Estimate catch per unit effort - `?estimate_harvest_rate` - Estimate harvest per unit effort - `?estimate_total_catch` - Estimate total catch - `?estimate_total_harvest` - Estimate total harvest - `?add_ages` - Attach age data to a design - `?est_age_distribution` - Estimate weighted age frequency - `?est_mean_age` - Compute design-weighted mean age - `?example_interviews` - Example interview dataset - `?example_ages` - Example age dataset For the effort estimation workflow, see the "Getting Started with tidycreel" vignette.