--- title: "Getting Started with tidycreel" output: rmarkdown::html_vignette: highlight: null vignette: > %\VignetteIndexEntry{Getting Started with tidycreel} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) ``` ## Introduction tidycreel gives fisheries biologists a tidy, pipe-friendly way to design and analyse creel surveys. You can work with the terms used in the field—dates, strata, counts, and effort—without having to work directly with the internals of the survey package. The basic workflow has three steps: 1. Define the survey calendar and its strata. 2. Attach count observations to the design. 3. Estimate effort and its variance. ## Survey Design We start by loading the package and the example calendar dataset: ```{r design} library(tidycreel) # Load example calendar data data(example_calendar) head(example_calendar) # Create a creel design with weekday/weekend strata design <- creel_design(example_calendar, date = date, strata = day_type) print(design) ``` The `creel_design()` function uses tidy selectors, so you can specify columns by name without quotes. The design object captures the survey structure: 14 days with weekday/weekend stratification. ## Adding Count Data Next, we attach the daily effort observations to the design. The `effort_hours` column holds angler-hours already accumulated over each day, not the raw angler count seen at a single moment — `estimate_effort()` expands whatever column it is given without converting units, so this design reports angler-hours: ```{r counts} # Load example count data data(example_counts) head(example_counts) # Attach counts to the design design <- add_counts(design, example_counts) print(design) ``` The `add_counts()` function validates that the count data matches the design structure, then constructs the internal survey design object. Notice that the design now shows count data attached with 14 observations. ## Estimating Total Effort With count data attached, we can estimate total effort across the entire survey period: ```{r total} # Estimate total effort result <- estimate_effort(design) print(result) ``` The result shows the estimated total effort, standard error, and 95% confidence interval. The estimate is 372.5 angler-hours over the 14-day period. Every calendar day was sampled here, so the season expansion factor is 1 and the estimate is exactly the sum of the attached `effort_hours`. ## Grouped Estimation We can also compute estimates separately for each stratum or group using the `by` parameter: ```{r grouped} # Estimate effort by day_type result_by_day <- estimate_effort(design, by = day_type) print(result_by_day) ``` The grouped results show separate estimates for weekday and weekend periods: approximately 202 angler-hours on weekends against 171 on weekdays. Compare the totals with care — the calendar holds 10 weekdays to 4 weekend days, so the modest gap between the totals reflects a much larger gap per day, roughly 50 angler-hours per weekend day against 17 per weekday. The `by` parameter accepts tidy selectors, so you can group by multiple columns or use tidyselect helpers like `starts_with()`. ## Variance Methods By default, `estimate_effort()` uses Taylor linearization for variance estimation. The package also supports bootstrap and jackknife methods: ```{r variance} # Bootstrap variance estimation (500 replicates) set.seed(123) # For reproducibility result_boot <- estimate_effort(design, variance = "bootstrap") print(result_boot) # Jackknife variance estimation result_jk <- estimate_effort(design, variance = "jackknife") print(result_jk) ``` **When to use each method:** - **Taylor linearization** (default): Computationally efficient and appropriate for most smooth statistics. This is the recommended default. - **Bootstrap**: Use when working with non-smooth statistics or when you want to verify Taylor linearization assumptions. More computationally intensive. - **Jackknife**: Alternative resampling method that is deterministic (unlike bootstrap). Useful for verification or when bootstrap is too slow. All three methods work with grouped estimation as well: ```{r grouped_variance} # Grouped estimation with bootstrap variance set.seed(123) result_grouped_boot <- estimate_effort(design, by = day_type, variance = "bootstrap") print(result_grouped_boot) ``` ## Schedule-Defined Special Strata If your survey calendar includes prospective high-use or other special periods, the same three-step workflow still applies. The difference is that the schedule may carry a resolved `final_stratum` column from `generate_schedule(..., special_periods = ...)`, and that resolved stratum should drive the analysis design. ```{r special-strata-overview, eval = FALSE} sched <- generate_schedule( start_date = "2027-07-24", end_date = "2027-08-04", n_periods = 1, sampling_rate = 0.5, include_all = TRUE, special_periods = opener_periods, seed = 42 ) calendar_for_design <- transform( sched[, c("date", "final_stratum")], analysis_stratum = ifelse(grepl("^high_use", final_stratum), final_stratum, "regular") )[, c("date", "analysis_stratum")] design_special <- creel_design( calendar_for_design, date = date, strata = analysis_stratum ) ``` Once counts and interviews are attached, `estimate_effort(..., target = "period_total")` and the total-catch/product estimators use the declared analysis strata directly. If one of those strata is too sparse for variance estimation, tidycreel names the sparse stratum in its diagnostic instead of failing opaquely. ## Next Steps This vignette covers the core tidycreel workflow for instantaneous count surveys. For more details on specific functions, see their help pages: - `?creel_design` - Define survey calendar and stratification - `?add_counts` - Attach count data to a design - `?estimate_effort` - Compute effort estimates with variance - `?as_creel_svydesign` - Extract internal survey object for advanced use For information on the example datasets: - `?example_calendar` - Example survey calendar - `?example_counts` - Example count observations