This vignette describes the syntax for specifying stochastic and
identity equations, priors, and lags in the koma
package.
Stochastic (regression) equations model a dependent variable with an error term. An intercept is included by default.
Identity equations enforce exact relationships.
# Identity equations with explicitly defined weights:
# To aggregate the component growth rates into a growth rate for GDP we need to define weights.
# This is done by specifying the weights in the equation.
# You can, e.g. use the nominal level weights of the last observed period.
gdp == 0.7*consumption + 0.2*investment + 0.2*government - 0.1*net_exports Lags are specified with L() or lag()
notation. Ranges and combinations are supported.
There are two kinds of priors in koma equations:
Coefficient priors are written in front of the term they belong to:
For example:
This sets a prior with mean 0.4 and variance
0.1 on the coefficient of gdp.
You can use priors on:
1 or
constantThe error-term prior is different. It is written as a final prior with no variable name:
In the error-term prior, the two values specify:
Some valid examples for priors are:
# Prior on the intercept
consumption ~ {0, 1000} 1 + gdp
# Prior on a lagged term
consumption ~ gdp + {0.9, 10} consumption.L(1)
# Same prior on several lags (applies to consumption.L(1), .L(2) and .L(4))
consumption ~ gdp + {0, 1000} consumption.L(1:2, 4)
# Prior on an endogenous regressor
consumption ~ {0.2, 0.5} service + gdp
# Error-term prior: {df, scale}
consumption ~ gdp + consumption.L(1) + {3, 0.001}Rules:
{mean, variance} variable.{df, scale} and
appears without a variable name.x.L(1:3),
x.L(1, 3), lag(x, 1:3)) applies to every lag
it expands to.You can override the default tau in your gibbs_settings
for a single equation by appending [tau = value] after its
equation. If the acceptance rate falls outside 30%-60 %, a warning is
emitted.
Dummy variables can be written as dummies(prefix, spec)
instead of spelling out every term. spec follows the same
syntax as lag ranges (a single index, a lower:upper range,
or a comma-separated mix).
# Shorthand:
consp ~ ydispbr + consp.L(1) + dummies(covid, 1:8)
# Equivalent to:
consp ~ ydispbr + consp.L(1) +
covid_1 + covid_2 + covid_3 + covid_4 + covid_5 + covid_6 + covid_7 + covid_8dummies() only expands the equation string - it
does not create any data. It is expanded before validation, so the
expanded names are treated exactly like any hand-typed variable, which
means the usual three steps for an exogenous variable still apply:
1. The data has to exist. Each expanded name
(covid_1, covid_2, …) must be a real 0/1
series in the ts_data passed to estimate().
koma does not generate this from a period specification -
you build it like any other series, e.g. as a shock in a single period
per dummy:
covid_periods <- c(2020.25, 2020.5, 2020.75, 2021, 2021.25, 2021.5, 2021.75, 2022)
for (i in seq_along(covid_periods)) {
dummy <- stats::ts(0, start = stats::start(ts_data$consp), end = stats::end(ts_data$consp),
frequency = stats::frequency(ts_data$consp))
window(dummy, start = covid_periods[i], end = covid_periods[i]) <- 1
ts_data[[paste0("covid_", i)]] <- as_ets(dummy, series_type = "rate", method = "none")
}2. It must be declared as exogenous, same as for any other regressor:
3. If forecasting, the dummy series must also extend
through the forecast horizon in ts_data (typically as 0,
since a one-off shock dummy shouldn’t recur) - forecast()
errors if an exogenous series doesn’t reach the forecast end date.