--- title: "Getting started with rbiogeme" author: "rbiogeme contributors" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting started with rbiogeme} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include = FALSE} knitr::opts_chunk$set( echo = TRUE, eval = FALSE, collapse = TRUE, comment = "#>" ) ``` `rbiogeme` lets an R user write a Biogeme model specification in R while the native Biogeme engine performs compilation, likelihood evaluation, differentiation, integration, optimization, simulation, and reporting. The R package does not create a second numerical engine and does not require the user to manipulate Python objects. This guide is intentionally self-contained. All expressions in the examples below are R expressions, and all names passed to the model are preserved when the specification is compiled by the bridge. The code chunks are shown as runnable examples but are not evaluated while the package documentation is built. Estimation and simulation require the user's configured Python environment and can create native output files. Any operation that creates persistent native files requires an explicit output directory; the package never silently writes those files to the current working directory. ## Installation and Python configuration Install a source tarball with ordinary R tools. When working from a checkout, build the tarball first with `R CMD build .` and install the resulting file. ```{r installation} install.packages( "/path/to/rbiogeme_0.1.2.tar.gz", repos = NULL, type = "source" ) ``` The package requires R 4.3 or later and Python 3.12 or later. The recommended path is to let `biogeme_setup()` provision the native requirement. If you manage Python yourself, pass the interpreter to `biogeme_setup()` before the first operation that initializes Python: ```{r configuration} library(rbiogeme) biogeme_setup( python = "/absolute/path/to/python", biogeme_requirement = "biogeme==3.3.5" ) # This reports the versions visible to reticulate. biogeme_diagnostics() ``` If you do not already manage Python, the shorter recommended path is: ```{r managed-setup} library(rbiogeme) check <- biogeme_setup() stopifnot(check$ready) ``` The default native requirement is `biogeme==3.3.5`. With no explicit interpreter, `biogeme_setup()` asks reticulate to provision an isolated managed environment and installs the requirement there on first use. With an explicit `python`, it selects that existing interpreter and verifies it; it does not install packages into a user-managed environment. Configuration is session-wide. Calling `biogeme_config()` or `biogeme_setup()` with a different interpreter after Python has already been initialized raises an error so that a model cannot accidentally use a different interpreter from the one selected. `biogeme_check()` is the recommended first troubleshooting step. It reports whether the configured runtime is ready and gives a corrective action for each failure. Use `biogeme_diagnostics()` when the detailed version list is needed. For a shorter first-run check after setup, use `biogeme_check()`. It catches runtime initialization failures, verifies the minimum R and Python versions, confirms that Biogeme can be imported, and gives a corrective action for each failure: ```{r readiness} check <- biogeme_check() if (!check$ready) { print(check) stop("The rbiogeme environment is not ready.") } ``` There are two supported setup paths. If you do not already manage Python, omit `python` and let `biogeme_setup()` provision `biogeme==3.3.5`. If you already have a Python environment, install Biogeme into it before starting R. For example, outside R: ```text python3.12 -m venv /path/to/rbiogeme-venv /path/to/rbiogeme-venv/bin/python -m pip install "biogeme==3.3.5" ``` Then select that interpreter before any operation that initializes Python: ```{r existing-environment} biogeme_setup(python = "/path/to/rbiogeme-venv/bin/python") ``` Selecting an existing interpreter does not install Biogeme into it. If configuration fails because Python has already been initialized, restart R and call `biogeme_config()` before constructing a database or model. ## Data and expressions Biogeme databases contain numeric columns. A database copies its input data at construction time, so later changes to the original data frame do not change the model specification. ```{r data} data <- data.frame( choice = c(1, 2, 1, 2, 1, 2, 1, 2), time = c(10, 8, 12, 7, 11, 9, 13, 8), cost = c(5, 7, 6, 8, 5, 7, 6, 9), income = c(1, 2, 1, 3, 2, 1, 3, 2) ) database <- biogeme_database("demo", data) # variable() is symbolic: it refers to a database column, not to an R vector. time <- variable("time") cost <- variable("cost") income <- variable("income") # biogeme_beta() creates a named native parameter. The name is part of the # equivalence contract and will appear unchanged in the result. b_time <- biogeme_beta("b_time", start = 0) b_cost <- biogeme_beta("b_cost", start = 0) asc_2 <- biogeme_beta("asc_2", start = 0) utility_1 <- b_time * time + b_cost * cost utility_2 <- asc_2 + b_time * time + b_cost * cost ``` Arithmetic operators build a neutral expression tree. They do not calculate a likelihood in R. Constants such as `0` are accepted wherever a scalar expression is expected. Comparisons and logical operators are symbolic too. They are useful for availability, filtering, piecewise definitions, and subsets: ```{r expression-syntax} available_2 <- (cost < 10) & (income >= 1) not_available_2 <- !(available_2) either_condition <- (time < 9) | (cost > 8) # Native-safe mathematical primitives are available as expression functions. safe_probability <- logzero(logit_probability( utilities = list(`1` = utility_1, `2` = utility_2), availability = list(`1` = 1, `2` = available_2), alternative = variable("choice") )) ``` Use `logzero()` for a native numerically safe logarithm. Other commonly used functions include `normal_cdf()`, `normal_pdf()`, `safe_exp()`, `sqrt()`, `abs()`, `biogeme_min()`, `biogeme_max()`, `Elem()`, `piecewise()`, `boxcox()`, and `derive()`. ## A complete multinomial logit model `logit_model()` is the shortest route for a cross-sectional multinomial logit model. The names of `utilities` identify alternatives. If those names are integer strings, they are also used as the observed choice codes. ```{r mnl-model} model <- logit_model( database = database, choice = "choice", utilities = list( `1` = utility_1, `2` = utility_2 ), availability = list( `1` = 1, `2` = available_2 ) ) # Check the native specification before running the optimizer. validation <- validate_model(model) stopifnot(validation$valid) # A fresh temporary output directory prevents an old YAML or iteration file # from being reused during an equivalence test. output_directory <- tempfile("rbiogeme-demo-") dir.create(output_directory) fit <- estimate( model, model_name = "rbiogeme_demo", control = biogeme_control( output_directory = output_directory, generate_html = FALSE, generate_yaml = FALSE, save_iterations = FALSE ) ) ``` `estimate()` always performs a fresh native estimation. The result is an R object containing serialized native results, not a live Python result object. The usual R methods expose the central post-estimation information: ```{r result-methods} summary(fit) coef(fit) vcov(fit) logLik(fit) nobs(fit) ``` The names in `coef(fit)` are the names supplied to `biogeme_beta()`. Do not rename or reorder them when comparing an R fit with a native fit. ## Choosing an output directory Native report and checkpoint files are opt-in and must have an explicit destination. For example, a user-facing run can choose a project directory: ```{r output-directory} output_directory <- "/absolute/path/to/my-biogeme-results" fit_with_reports <- estimate( model, model_name = "my_model_with_reports", control = biogeme_control( output_directory = output_directory, generate_html = TRUE, generate_yaml = TRUE, save_iterations = TRUE ) ) ``` Use `tempdir()` instead when the files are only intermediate artifacts in a test or a short demonstration. The command-line examples in `inst/examples/` follow the same rule through their required `--output=/path/to/output` argument. Before estimation, `validate_model(model)` performs native specification validation without running the optimizer or writing estimation results. After fitting a logit model, `predict(fit)` evaluates one native probability column per alternative. Scenario predictions use `predict(fit, newdata = data.frame(...))`; the supplied data frame must contain the variables referenced by the model. ## Generic likelihoods and simulations Specialized constructors are convenient, but every model can be expressed with `biogeme_model()`. This is the general interface for a complete likelihood, an optional probability, named simulation expressions, weights, panels, draws, subsets, and parameter overrides. ```{r generic-model} probability <- logit_probability( utilities = list(`1` = utility_1, `2` = utility_2), availability = list(`1` = 1, `2` = available_2), alternative = variable("choice") ) generic_model <- biogeme_model( database = database, formula = logzero(probability), probability = probability, simulations = list( probability = probability, time_cost_ratio = time / cost ) ) generic_fit <- estimate( generic_model, model_name = "rbiogeme_generic", control = biogeme_control( output_directory = output_directory, generate_html = FALSE, generate_yaml = FALSE, save_iterations = FALSE ) ) simulated <- simulate( generic_model, beta = generic_fit, control = biogeme_control(output_directory = output_directory) ) as.data.frame(simulated) ``` The `formula` is the expression used for estimation. `probability` documents the corresponding probability when a generic model has one, while `simulations` is a named list of expressions evaluated by native Biogeme at the supplied estimates. ## Database operations Database transformations can remain symbolic until the bridge compiles the complete model. This keeps derived-variable and filtering semantics in native Biogeme and avoids an R callback during numerical evaluation. ```{r database-operations} database_with_ratio <- biogeme_database_define_variable( database, name = "time_cost_ratio", expression = variable("time") / variable("cost") ) database_without_high_cost <- biogeme_database_remove( database_with_ratio, condition = variable("cost") > 8 ) biogeme_database_columns(database_without_high_cost) biogeme_database_nrow(database_without_high_cost) biogeme_database_filtered_row_count(database_without_high_cost) biogeme_database_row_ids(database_without_high_cost) # Native derived columns and filters are materialized explicitly when their # resulting data frame or row count is needed in R. materialized <- biogeme_database_materialize(database_without_high_cost) as.data.frame(materialized) ``` For panels, the panel identifier must be a numeric column and observations for each individual must already be contiguous. The constructor validates that ordering before the model is compiled: ```{r panel-database} panel_data <- data.frame( person = c(1, 1, 2, 2, 3, 3), choice = c(1, 2, 2, 1, 1, 2), time = c(10, 8, 9, 11, 12, 7) ) panel_database <- biogeme_panel_database( name = "demo_panel", data = panel_data, panel_id = "person" ) biogeme_database_is_panel(panel_database) ``` ## Reproducible output Native estimation can create YAML, HTML, iteration, NetCDF, checkpoint, or diagnostic files depending on the operation. Use an operation-specific temporary directory in tests and examples. Disable each file type that is not part of the behavior under test: ```{r reproducibility} clean_control <- biogeme_control( output_directory = tempfile("rbiogeme-output-"), seed = 1234, generate_html = FALSE, generate_yaml = FALSE, save_iterations = FALSE ) ``` For operations where loading an old result is the feature being demonstrated, use `estimate_or_load()` explicitly and set `force` deliberately. Ordinary `estimate()` calls do not silently recycle a previous result. ## Finding the right starting point | Goal | Start with | | --- | --- | | Estimate a standard choice model | `logit_model()` and `estimate()` | | Specify a custom likelihood | `biogeme_model()` | | Work with panel observations | `biogeme_panel_database()` and `panel_likelihood_trajectory()` | | Evaluate probabilities or scenarios | `predict()` and `simulate()` | | Use Bayesian, MDCEV, Monte Carlo, catalog, hybrid-choice, or sampling features | The corresponding specialized constructor and `advanced-models` | The reference manual is available through R's normal help system. The most useful entry points are: ```{r help} ?rbiogeme ?biogeme_check ?biogeme_model ?estimate ?simulate ``` If a first model does not run, use this order: 1. Run `biogeme_check()` and resolve every `ERROR` row. 2. Run `biogeme_database_columns(database)` and compare the result with the names used by `variable()`. 3. Run `validate_model(model)` to check the compiled native specification without running the optimizer. 4. Use `biogeme_config(debug = TRUE)` in a fresh R session if the bridge error still needs its native traceback. For more specialized model families, continue with `modeling-workflows` and `advanced-models`.