---
title: "Getting started with zuhtml"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Getting started with zuhtml}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```
zuhtml turns real-world HTML into ordinary R values: character vectors,
lists and data frames. It parses the way a browser does, so malformed
markup is repaired rather than rejected. You give it a string, raw bytes,
a file, a URL or a connection.
```{r}
library(zuhtml)
```
## A page to work with
A small catalogue page, as it might have been saved from a site. Some
markup is sloppy on purpose: unclosed `
`s, unquoted attributes, a
stray end tag, and a product card without a price.
```{r}
page <- '
Tea shop
'
doc <- html_parse(page, base_url = "https://example.org/shop/")
doc
```
`html_read()` does the same for a file. `base_url` is where the page came
from; relative links are resolved against it.
## Selecting elements
`html_elements()` finds every element that matches a CSS selector.
`html_element()` finds the first match below *each* input node, and keeps
a missing node where there is none. That is what keeps extracted columns
aligned when some records lack a field:
```{r}
cards <- html_elements(doc, ".product")
cards
products <- data.frame(
name = html_text_clean(html_element(cards, ".name")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
products
```
Genmaicha has no price, so it gets `NA` rather than shifting the column.
## Values
`html_text_clean()` gives text as a reader wants it; `html_text()` gives
it exactly as parsed. `html_attr()` reads attributes, and
`html_serialize()` writes nodes back as HTML.
```{r}
html_text_clean(html_element(doc, "title"))
html_attr(html_elements(doc, "nav a"), "href")
html_serialize(html_element(doc, "h2"))
```
## Structures
Links, lists and tables have their own extractors:
```{r}
html_links(doc, absolute = TRUE)
lapply(html_elements(doc, "nav ul"), html_list)
html_tables(doc)
```
Table columns are character: `"0100"` keeps its leading zero. Convert
types yourself when you know them, for example with `type.convert()`.
## What the parser repaired
Real pages nearly always have markup errors, which the parser repairs.
`html_problems()` lists them:
```{r}
html_problems(doc)
```
## Where next
* `vignette("selectors")`: the supported CSS subset.
* `vignette("tables-and-lists")`: how tables and lists are read.
* `vignette("limits-and-encoding")`: resource limits, encodings, and what
zuhtml does not do.