--- title: "Reading dlf files" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Reading dlf files} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} # nolint start: indentation-linter knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) # nolint end ``` ```{r setup} library(daisytools) ``` Daisy log files (.dlf) can be read with `read_dlf`. `read_dlf` can read a single dlf file, a directory tree of dlf files produced by the daisy spawn program, or a directory tree of dlf files in any order. When called with just a path the function will automatically guess how it should process the path. You can also explicitly control how the path is processed by setting the mode parameter. ## Reading a single dlf file We start by reading a single dlf file and inspect the contents. We use the Soil chemical log that is bundled with `daisytools`. ```{r} data_dir <- system.file("extdata", package="daisytools") path <- file.path(data_dir, "hourly/P2D-Daily-Soil_Chemical_110cm.dlf") dlf <- read_dlf(path) head(dlf) ``` `read_dlf` returns an S4 object of class `Dlf`. The data in the dlf file is stored in three _slots_ in the `Dlf` object, `header`, `units`, `data` ```{r} slotNames(dlf) ``` The `header` slot is a list of the metadata that is stored before the actual data in the dlf file. The `units` slot is a `data.frame` with unit information for each data column. The `data` slot is a `data.frame` containing the data. A slot is accessed with the `@` operator ```{r} dlf@header ``` ```{r} dlf@units ``` ```{r} head(dlf@data) ``` We can also directly access and assign columns of the `data` slot with the `$` and `[[` operators ```{r} head(dlf$year) dlf$year <- dlf$year + 2 head(dlf[["year"]]) dlf[[1, "year"]] <- dlf[[1, "year"]] + 10 dlf[[1, "year"]] ``` ## Reading a depth distributed output Some log files contain the distribution of a variable over depth and time. By default, these are detected and converted automatically by `read_dlf` ```{r} data_dir <- system.file("extdata", package="daisytools") path <- file.path(data_dir, "daily/DailyP/DailyP-Daily-WaterFlux.dlf") dlf <- read_dlf(path) head(dlf) ``` It is possible to disable this conversion ```{r} dlf <- read_dlf(path, convert_depth=FALSE) head(dlf@data) ``` ## Reading Daisy spawn output Daisy spawn produces a directory hierarchy where each directory contains the same log files. We use the Daisy spawn outputs that are bundled with `daisytools`. ```{r} data_dir <- system.file("extdata", package="daisytools") path <- file.path(data_dir, "daisy-spawn-like") list.files(path, recursive=TRUE) ``` Notice that each directory contains the same log files. ```{r} dlfs <- read_dlf(path) ``` Instead of a single Dlf object, we now get a named list of Dlf objects ```{r} typeof(dlfs) dlf_names <- names(dlfs) dlf_names ``` The list names correspond to the log types that was read. All logs of the same type are concatenated together, and an extra column is added to the data field. The name of the column is "sim" by default, but can be set by passing `col_name="MySimulationName"` to `read_dlf`. ```{r} name <- dlf_names[1] name head(dlfs[[name]]@data) unique(dlfs[[name]]$sim) ``` ## Reading a directory with dlf files We use the tracer and field nitrogen logs that are bundled with `daisytools` and read them with `read_dlf_dir` ```{r} data_dir <- system.file("extdata", package="daisytools") path <- file.path(data_dir, "annual") list.files(path, pattern=".*dlf", recursive=TRUE) ``` Notice that the directories contains a different set of log files. ```{r} dlfs <- read_dlf(path) ``` Instead of a single Dlf object, we now get a named list of Dlf objects ```{r} names(dlfs) ``` The list of names correspond to file paths, which can be a bit unwieldy. We can either manually change the names or use the convenience functions `drop_dir_from_names` and `strip_common_prefix_from_names` ```{r} dlfs <- drop_dir_from_names(dlfs) names(dlfs) dlfs <- strip_common_prefix_from_names(dlfs) names(dlfs) ``` ```{r} name <- names(dlfs)[1] name dlfs[[name]]@header ``` ### Controlling which files are included `read_dlf` takes an optional pattern parameter that is used for modes "spawn" and "dir" and controls which files are included. By default all files with the suffix `.dlf` are included. If we only want to read files ending with `2b.dlf`, we can do ```{r} data_dir <- system.file("extdata", package="daisytools") path <- file.path(data_dir, "annual") dlfs <- read_dlf(path, pattern=".*2b\\.dlf") names(dlfs) ``` `pattern` is a regular expression. You should around with it and see if you can find a shorter pattern that results in the same files being read.