--- title: "Getting started with danlex" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting started with danlex} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- danlex provides access to Danish legislation published by Retsinformation, the official state legal information system. This vignette works through a single instrument in detail and then scales up to the corpus. ``` r library(danlex) ``` This vignette was pre-computed: the code was run against the live service when the package was built, so the results below are real but the examples do not call the API when the package is checked or installed. ## Finding a document Danish acts and regulations are cited by year and number — "bekendtgørelse nr. 763 af 1998". That maps directly onto an identifier: ``` r doc <- dlx_get_doc("lta", 1998, 763) doc$title #> [1] "Bekendtgørelse om indbetaling af Den Særlige Pensionsopsparing for medlemmer af arbejdsløshedskasser for lønmodtagere ved udbetaling af dagpenge m.v." doc$document_type #> [1] "Bekendtgørelse" doc$ministry #> [1] "Beskæftigelses- og Ligestillingsministeriet" ``` The first argument is the collection. `lta` is *Lovtidende A*, which carries both acts and regulations in a single number space — so the instrument above is a *bekendtgørelse* even though it sits in the same series as acts. The type is never inferable from the identifier alone, which is why `document_type` is always returned. The other collections are `ltb` and `ltc` (*Lovtidende B* and *C*, the latter carrying treaties), `mt` (*Ministerialtidende*, discontinued in 2012), `retsinfo` (administrative decisions), `fob` (Ombudsman decisions) and `ft` (parliamentary documents). A document that does not exist is a normal answer, not an error: ``` r dlx_get_doc("lta", 1998, 999999)$status #> [1] "not_found" ``` ## Text availability Retsinformation returns full text only for material published from late September 2007 onward. Earlier documents carry metadata alone, and `status` reports this: ``` r doc$status #> [1] "metadata_only" ``` The boundary is sharp. Act 1080 of 2007, announced on 14 September, has no text; act 1081, announced on 25 September, does: ``` r dlx_get_doc("lta", 2007, 1080)$status #> [1] "metadata_only" dlx_get_doc("lta", 2007, 1081)$status #> [1] "ok" ``` This is a property of the source. Metadata for earlier material is complete and reaches back to 1852 for Lovtidende A — and, in the `retsinfo` collection, to 1665. ## Reading a document For a modern instrument, `dlx_get_text()` returns the running text with its structural markers in place: ``` r txt <- dlx_get_text("lta", 2025, 50) txt$n_char #> [1] 9601 cat(substr(txt$text, 1, 400)) #> I medfør af § 7 b, stk. 2, § 22, stk. 1, nr. 1, 3, 4, 5, 6 og 7, og stk. 3-4, § 23 og § 30 i lov om erhvervsakademiuddannelser og professionsbacheloruddannelser, jf. lovbekendtgørelse nr. 396 af 12. april 2024, fastsættes efter bemyndigelse i henhold til § 4, stk. 1, i bekendtgørelse nr. 1118 af 28. oktober 2024 om delegation af uddannelses- og forskningsministerens beføjelser til Uddannelses- og ``` ## Working with provisions Danish legal drafting is highly structured, and Retsinformation preserves that structure in its XML. `dlx_get_paragraphs()` returns one row per subsection, so provisions can be handled as data rather than as text to be parsed: ``` r p <- dlx_get_paragraphs("lta", 2025, 50) nrow(p) #> [1] 25 p[1:4, c("paragraf", "stk", "text")] #> # A tibble: 4 × 3 #> paragraf stk text #> #> 1 § 1. Formålet med erhvervsakademiuddannelsen som maritim teknolog… #> 2 § 1. Stk. 2. Den uddannede skal have viden, færdigheder og kompetencer (l… #> 3 § 1. Stk. 3. Den uddannede, som har gennemført de maritime valgfag, skal … #> 4 § 2. Uddannelsen giver den uddannede ret til at anvende titlen ma… ``` Note that the first subsection of each section carries no marker: `Stk. 1.` is implicit in Danish drafting practice and simply absent from the source. ### Amending instruments An amending instrument contains provisions that are not its own — they are new text destined for another instrument. The `is_amendment` column makes the distinction: ``` r own <- dlx_get_paragraphs("lta", 2025, 50) amd <- dlx_get_paragraphs("lta", 2025, 1) table(own$is_amendment) #> #> FALSE #> 25 table(amd$is_amendment) #> #> TRUE #> 12 ``` The second instrument amends a regulation from 2018; every provision it contains belongs to that regulation, not to itself. Treating the two alike would badly distort any corpus-level analysis of what Danish law says. ## References between instruments Regulations record the provisions they were issued under, and `dlx_get_references()` returns these as an edge table: ``` r refs <- dlx_get_references("lta", 1998, 763) refs[, c("ref_accn", "ref_date", "ref_title")] #> # A tibble: 2 × 3 #> ref_accn ref_date ref_title #> #> 1 A19990059229 1999-07-14 Bekendtgørelse af lov om arbejdsløshedsforsikring m.v. #> 2 A19990066629 1999-08-19 Bekendtgørelse af lov om en aktiv arbejdsmarkedspolit… ``` The `ref_accn` values are accession numbers, which `dlx_get_doc()` accepts directly — so a citation network can be traversed without building a lookup table first: ``` r parent <- dlx_get_doc(accn = refs$ref_accn[1]) parent$title #> [1] "Bekendtgørelse af lov om arbejdsløshedsforsikring m.v." parent$document_type #> [1] "Lovbekendtgørelse" ``` ### References are maintained, not historical One property deserves emphasis before anyone builds a network from this data. The referring instrument is from 1998, yet its references are dated 1999: ``` r refs$ref_date #> [1] "1999-07-14" "1999-08-19" ``` This is not an error. Retsinformation updates these links to point at the *current* consolidated version of the parent act, so they reflect the state of the law now rather than at the moment of enactment. An analysis that treats the edges as timestamped to the referring document's date will be systematically wrong. ## Scaling up ### Enumerating the corpus `dlx_list_documents()` retrieves the full index of everything published through the service: ``` r idx <- dlx_list_documents(progress = FALSE) nrow(idx) #> [1] 202924 table(idx$collection) #> #> fob ft lta ltb ltc mt retsinfo #> 2957 41492 63218 45 4783 6438 83991 ``` The first call takes about a minute, because the index is published as a sitemap split across 21 pages whose ordering does not permit fetching a single year. The result is held for the rest of the session, so later calls are immediate. Nothing is written to disk. To keep the index between sessions, save it: ``` r saveRDS(idx, "danlex-index.rds") ``` The index carries identifiers and modification dates only. Note that `lastmod` records when Retsinformation last revised its *record*, which is unrelated to when the instrument was enacted or repealed. ### Coverage varies by collection The index is the quickest way to see what is actually available: ``` r lta <- subset(idx, collection == "lta" & !is.na(year)) range(lta$year) #> [1] 1852 2026 mt <- subset(idx, collection == "mt" & !is.na(year)) max(mt$year) #> [1] 2012 ``` Ministerialtidende ends in 2012 because it was discontinued. Ombudsman decisions (`fob`) begin in 1980, and treaties (`ltc`) drop off sharply after 2008. These are genuine historical discontinuities rather than gaps in coverage, and an empty result for a given year may simply mean nothing was published. ### Retrieving in bulk `dlx_get_docs()` takes a vector of identifiers. Here is every treaty published in 2025: ``` r treaties <- dlx_list_documents(year = 2025, collection = "ltc", progress = FALSE) docs <- dlx_get_docs(treaties$eli, progress = FALSE) nrow(docs) #> [1] 6 table(docs$status) #> #> ok #> 6 ``` Failures are returned as rows with `status = "error"` rather than raised as conditions, so a long run survives a dropped connection: ``` r sum(docs$status == "error") #> [1] 0 ``` Once retrieved, the metadata is ordinary data: ``` r sort(table(docs$ministry), decreasing = TRUE) #> #> Udenrigsministeriet Skatte- og Vækstministeriet #> 5 1 ``` ## Tracking changes The index is refreshed monthly, so recent material may be missing from it. `dlx_get_changes()` reads a separate feed covering at least the last 60 days: ``` r ch <- dlx_get_changes() nrow(ch) #> [1] 405 range(ch$updated) #> [1] "2026-07-02" "2026-08-29" table(ch$published_in) #> #> Lovtidende A Lovtidende C Retsinformation #> 101 2 302 ``` The feed itself records only the publication channel, not the collection. danlex derives `collection` from the identifier, which makes it possible to separate legislation from administrative decisions: ``` r recent_acts <- dlx_get_changes(collection = "lta", reason = "NewDocument") nrow(recent_acts) #> [1] 83 head(recent_acts$title, 3) #> [1] "Bekendtgørelse om Ledningsejerregistret" #> [2] "Bekendtgørelse om landsplandirektiv for to arealer til højspændingsstationer i hhv. Køge Kommune og Stevns Kommune" #> [3] "Bekendtgørelse om kriterier for vurdering af kommunale projekter vedrørende vandløbsrestaurering" ``` Because the feed is a static file with no query interface, these filters apply after retrieval. Asking for a date earlier than the feed reaches will return a truncated result and warn. ## Where to go next The whole corpus is around 200,000 documents, of which roughly 74,000 are legislation in the strict sense. Full text is available from late 2007 onward, structured down to the subsection, with maintained cross-references throughout. That combination supports work that is awkward against most national legal sources: provision-level text analysis, citation network construction, and long-run measurement of legislative output. danlex is part of **lexverse**, a family of R packages for legal and regulatory data alongside [eurlex](https://cran.r-project.org/package=eurlex), [finlex](https://cran.r-project.org/package=finlex) and swelex. ## Terms of use Retsinformation's data is provided by Civilstyrelsen. Consult the terms at before redistributing it or using it commercially.