--- title: "Limits, encodings and safety" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Limits, encodings and safety} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` ```{r} library(zuhtml) ``` ## Every call runs under limits HTML from the web is untrusted input. zuhtml bounds the work and memory any page can cost, per call, with `html_limits()`: ```{r} html_limits() ``` * `max_input` is checked before parsing. `max_memory` bounds the parser's native memory and the document it builds; it is the real guard, since some markup costs far more memory per byte than other markup. * `max_depth` bounds nesting *while parsing*. Tree construction takes time quadratic in nesting depth, so a check after parsing would come too late: 100,000 nested elements would take 16 seconds. With the limit it fails at once. * `max_nodes`, `max_errors` (parse problems kept), `max_table_cells` and `max_selector_length` bound the rest. Exceeding a limit is a `zuhtml_limit_error`, raised after every native allocation has been released. It says which limit and by how much: ```{r} err <- tryCatch( html_parse(strrep("
Small page
", limits = strict) ``` ## Encodings A string is already text: it is used as UTF-8. Raw bytes are decoded, in order of preference, with a byte-order mark, the `encoding` you give, the page's own `` declaration, or UTF-8: ```{r} bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9)) # "caf\xe9" html_text_clean(html_parse(bytes, encoding = "latin1")) ``` Invalid input is an error, never silently replaced: ```{r} try(html_parse(bytes)) ``` The declaration is found as a browser finds it, by scanning the first 1024 bytes for `` or its `http-equiv` form. Labels mean what they mean to browsers, so `iso-8859-1` is read as windows-1252, which makes byte 0x93 a curly quote rather than a control character: ```{r} page <- c(charToRaw("
"), as.raw(0x93), charToRaw("Quoted"), as.raw(0x94)) doc <- html_parse(page) html_text_clean(doc) html_info(doc)[c("encoding", "encoding_source")] ``` When you fetch a page, pass the charset from the HTTP `Content-Type` header as `encoding`: it takes precedence over the page's declaration, as it does in a browser. A byte-order mark takes precedence over both; one that contradicts `encoding` is an error. ## Errors are classed Handle errors by class, never by message text: ```{r} tryCatch( html_elements(html_parse("
"), "p:hover"), zuhtml_selector_error = function(e) paste("unsupported at", e$position) ) ``` See `?zuhtml-conditions` for the classes and their fields. ## What zuhtml does not do * **It does not sanitize.** Parsing and serializing keep `