Equation Syntax Reference

Overview

This vignette describes the syntax for specifying stochastic and identity equations, priors, and lags in the koma package.

1. Stochastic Equations

Stochastic (regression) equations model a dependent variable with an error term. An intercept is included by default.

# With default intercept:
consumption ~ gdp + consumption.L(1)

# Without intercept:
consumption ~ gdp + consumption.L(1) - 1

# Explicit intercept:
consumption ~ 1 + gdp + consumption.L(1)

2. Identity Equations

Identity equations enforce exact relationships.

# Identity equations with explicitly defined weights:
# To aggregate the component growth rates into a growth rate for GDP we need to define weights.
# This is done by specifying the weights in the equation. 
# You can, e.g. use the nominal level weights of the last observed period.
gdp == 0.7*consumption + 0.2*investment + 0.2*government - 0.1*net_exports 

3. Injected Parameters

# Ratios computed from data:
gdp == (nom_cons/nom_gdp) * cons

4. Lag Notation

Lags are specified with L() or lag() notation. Ranges and combinations are supported.

# Single lag:
x.L(1)
lag(x, 1)

# Range of lags:
x.L(1:4)
lag(x, 1:4)

# Mix range and specific lags:
x.L(1:3, 5)

5. Priors

There are two kinds of priors in koma equations:

Coefficient priors are written in front of the term they belong to:

{mean, variance} variable

For example:

consumption ~ {0.4, 0.1} gdp + consumption.L(1)

This sets a prior with mean 0.4 and variance 0.1 on the coefficient of gdp.

You can use priors on:

consumption ~
  {0, 1000} 1 +
  {0.4, 0.1} gdp +
  {0.9, 10} consumption.L(1) +
  {0.2, 0.5} service

The error-term prior is different. It is written as a final prior with no variable name:

consumption ~ gdp + consumption.L(1) + {3, 0.001}

In the error-term prior, the two values specify:

Some valid examples for priors are:

# Prior on the intercept
consumption ~ {0, 1000} 1 + gdp

# Prior on a lagged term
consumption ~ gdp + {0.9, 10} consumption.L(1)

# Same prior on several lags (applies to consumption.L(1), .L(2) and .L(4))
consumption ~ gdp + {0, 1000} consumption.L(1:2, 4)

# Prior on an endogenous regressor
consumption ~ {0.2, 0.5} service + gdp

# Error-term prior: {df, scale}
consumption ~ gdp + consumption.L(1) + {3, 0.001}

Rules:

6. Equation-specific Tau

You can override the default tau in your gibbs_settings for a single equation by appending [tau = value] after its equation. If the acceptance rate falls outside 30%-60 %, a warning is emitted.

"consumption ~ constant + gdp + consumption.L(1) + consumption.L(2) [tau = 1.2]"

7. Dummy Variables

Dummy variables can be written as dummies(prefix, spec) instead of spelling out every term. spec follows the same syntax as lag ranges (a single index, a lower:upper range, or a comma-separated mix).

# Shorthand:
consp ~ ydispbr + consp.L(1) + dummies(covid, 1:8)

# Equivalent to:
consp ~ ydispbr + consp.L(1) +
  covid_1 + covid_2 + covid_3 + covid_4 + covid_5 + covid_6 + covid_7 + covid_8

dummies() only expands the equation string - it does not create any data. It is expanded before validation, so the expanded names are treated exactly like any hand-typed variable, which means the usual three steps for an exogenous variable still apply:

1. The data has to exist. Each expanded name (covid_1, covid_2, …) must be a real 0/1 series in the ts_data passed to estimate(). koma does not generate this from a period specification - you build it like any other series, e.g. as a shock in a single period per dummy:

covid_periods <- c(2020.25, 2020.5, 2020.75, 2021, 2021.25, 2021.5, 2021.75, 2022)
for (i in seq_along(covid_periods)) {
  dummy <- stats::ts(0, start = stats::start(ts_data$consp), end = stats::end(ts_data$consp),
    frequency = stats::frequency(ts_data$consp))
  window(dummy, start = covid_periods[i], end = covid_periods[i]) <- 1
  ts_data[[paste0("covid_", i)]] <- as_ets(dummy, series_type = "rate", method = "none")
}

2. It must be declared as exogenous, same as for any other regressor:

exogenous_variables <- c("ydispbr", paste0("covid_", 1:8))

3. If forecasting, the dummy series must also extend through the forecast horizon in ts_data (typically as 0, since a one-off shock dummy shouldn’t recur) - forecast() errors if an exogenous series doesn’t reach the forecast end date.