--- title: "Get started with tidymedia" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Get started with tidymedia} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) # Work in a temporary directory: the chunks below write real files, and none of # them belongs beside the vignette sources (verification.Rmd:33 set this). knitr::opts_knit$set(root.dir = tempdir()) ``` ```{r setup} library(tidymedia) ``` tidymedia runs [FFmpeg](https://ffmpeg.org/) and [MediaInfo](https://mediaarea.net/en/MediaInfo) from R. It helps you prepare media files for research in a way you can repeat. It trims, crops and converts files, often many at once. It also reads media metadata into tibbles. tidymedia does not try to cover everything FFmpeg can do. The words that FFmpeg uses, such as [codec](#glossary) and [stream](#glossary), are defined in the [glossary](#glossary) at the end of this page. This page uses a short sample clip that comes with the package: ```{r} video <- system.file("extdata", "sample.mp4", package = "tidymedia") ``` ## Start with a task function Most jobs need one call to a task function. For example, you may need the audio of a recording for a transcription tool. `extract_audio()` writes the audio to its own file: ```{r, eval = nzchar(Sys.which("ffmpeg")) && nzchar(Sys.which("ffprobe"))} extract_audio(video, "audio.m4a") ``` A task function runs FFmpeg at once. It returns the FFmpeg command it ran, but invisibly, so R prints nothing. A function that writes two files, such as `separate_audio_video()`, returns both commands. To see the command without running it, add `run = FALSE`. The function then returns the command as a string that you can read, log or save: ```{r} extract_audio(video, "audio.m4a", run = FALSE) ``` This is the main idea of the package. You can read each command before you run it. Cropping works the same way: ```{r} crop_video(video, "cropped.mp4", width = 160, height = 120, run = FALSE) ``` ### Choosing an audio track Some recordings have more than one audio track, for example a room microphone and a lapel microphone. If you do not choose a track, FFmpeg chooses one for you. Each task function that reads one input file and picks an audio track has an `audio_stream` argument. It counts from 0, and it counts only the audio tracks. So `audio_stream = 1` is the second audio track, wherever it sits in the file: ```{r} extract_audio(video, "lapel.m4a", audio_stream = 1, run = FALSE) ``` If you leave `audio_stream` out, the default depends on the function. A function that writes exactly one audio track, such as `extract_audio()`, takes the first track. A function that passes the audio through, such as `crop_video()`, keeps every track. Compare the `-map` parts of the two commands above and below: ```{r} crop_video(video, "cropped.mp4", width = 160, height = 120, run = FALSE) ``` The help page `?audio_stream` lists which functions use each default. It also explains `audio_input`, which the functions for several input files use. That argument counts input files from 0, not audio tracks. Each task function has a batch version for a folder of files, such as `extract_audio_batch()` and `crop_video_batch()`. See `vignette("batch")`. For a full research example that uses many task functions, see `vignette("workflow")`. ## Three kinds of function tidymedia has three kinds of function: - **Task functions**, such as `extract_audio()` and `crop_video()`, do one common job in one call. Start here. - **Pipeline functions**, whose names start with `ffm_`, build an FFmpeg command one step at a time. Each task function uses them. Use them when no task function does the job you need. - **Direct commands**, `ffmpeg()`, `ffprobe()` and `mediainfo()`, pass your own arguments to the program. Use them for anything that tidymedia does not cover. The rest of this page shows the pipeline functions. ## Building a pipeline A pipeline starts with `ffm_files()`, which names the input and output files. You add steps with `|>`. Each step adds an instruction, and nothing runs yet. `ffm_compile()` turns the pipeline into the FFmpeg command: ```{r} ffm_files(video, "output.mp4") |> ffm_trim(start = 1, end = 5) |> ffm_crop(width = 160, height = 120) |> ffm_codec(video = "libx264") |> ffm_compile() ``` When you print a pipeline, R shows the same command. So you can look at a pipeline at any point: ```{r} ffm_files(video, "output.mp4") |> ffm_scale(width = 320, height = 240) |> ffm_pixel_format("yuv420p") ``` To run the command and write the output file, use `ffm_run()` in place of `ffm_compile()`. ### More pipeline steps `ffm_fps()` changes the [frame rate](#glossary). `ffm_drawbox()` draws a box on the picture, which can hide a name on screen: ```{r} ffm_files(video, "boxed.mp4") |> ffm_fps(15) |> ffm_drawbox(x = 10, y = 10, width = 60, height = 40, color = "black") |> ffm_compile() ``` `ffm_loudnorm()` makes audio a set loudness, in [LUFS](#glossary), with a limit on its [true peak](#glossary). Here `ffm_drop()` also leaves the video out of the output: ```{r} ffm_files(video, "speech.m4a") |> ffm_drop("video") |> ffm_loudnorm(target_loudness = -23, true_peak = -1) |> ffm_compile() ``` `ffm_output_options()` adds FFmpeg output options that have no pipeline function of their own. tidymedia still puts them in the right place in the command: ```{r} ffm_files(video, "web.mp4") |> ffm_output_options("-movflags +faststart") |> ffm_compile() ``` ## Fast cuts and exact cuts You can cut a clip in two ways. An exact cut [re-encodes](#glossary) the video, which is slower. A fast cut uses a [stream copy](#glossary), which keeps the quality but starts at the nearest [keyframe](#glossary). `ffm_seek()` does both. Set `reencode = FALSE` and add `ffm_copy()` for a fast cut: ```{r} # Fast cut with no loss of quality ffm_files(video, "output.mp4") |> ffm_seek(start = 1, end = 5, reencode = FALSE) |> ffm_copy() |> ffm_compile() ``` `ffm_seek()` uses FFmpeg's `-ss` and `-to` options. `ffm_trim()` uses FFmpeg's `trim` filter. Only `ffm_seek()` can make a fast cut. ## Combining multiple inputs Some pipeline functions take more than one input. Give `ffm_files()` a vector of files. Then use `ffm_hstack()` to put the videos side by side, or `ffm_vstack()` to put one above the other. `ffm_overlay()` puts one video on top of another, and `ffm_concat()` joins them end to end: ```{r} ffm_files(c(video, video), "side_by_side.mp4") |> ffm_hstack() |> ffm_compile() ``` `ffm_hstack()`, `ffm_vstack()` and `ffm_overlay()` leave the audio out. To keep it, add `ffm_map("0:a")`. `ffm_concat()` keeps all the streams, audio included. One-input task functions that pass audio through, such as `crop_video()`, do the opposite and keep every audio track. Two task functions cover the common cases. `compare_videos()` puts videos side by side or one above the other. `picture_in_picture()` puts a smaller copy of one video on top of another. Both take `audio_input`, which names the input whose audio to keep. They copy that audio unchanged unless you choose an audio [encoder](#glossary) with `audio_codec`. ```{r} compare_videos(c(video, video), "compare.mp4", audio_input = 0, run = FALSE) ``` A pipeline has one input chain, a list of filters in order, and one output. It cannot build an FFmpeg filter graph with branches. For that, use the direct command `ffmpeg()`: ```{r, eval = nzchar(Sys.which("ffmpeg"))} # Your own arguments, passed to FFmpeg as they are ffmpeg("-version")[1] ``` ## Glossary - **Codec**: a way to compress audio or video, such as H.264 or AAC. FFmpeg names each codec with a short string, such as `"h264"` or `"aac"`. - **Container**: the file format that holds the streams, such as MP4, MKV or WAV. The file extension usually names the container. - **Stream**: one track in a file, such as the video or one audio track. A file can hold more than one stream of each kind. - **Encoder**: the part of FFmpeg that writes a stream in a codec. One codec can have several encoders, such as `libx264` and `h264_videotoolbox` for H.264. - **Re-encode**: to decode a stream and write it again with an encoder. This is slower and can lose some quality, but it lets FFmpeg change the picture or sound. - **Stream copy**: to put a stream into the output file without decoding it. This is fast and keeps the quality, but it cannot change the picture or sound. - **Pixel format**: how a video stores the color of each pixel, such as `"yuv420p"`. Many players need `"yuv420p"` to play H.264 video. - **Keyframe**: a video frame that is stored whole, not as a change from the frames before it. A stream copy can start a cut only at a keyframe. - **Frame rate**: the number of video frames per second. - **Sample rate**: the number of audio samples per second, in hertz, such as 48000. - **LUFS**: a unit of loudness that follows how loud people hear the sound. The EBU R 128 broadcast standard uses -23 LUFS. - **True peak**: the highest level the sound wave reaches, including between samples. A limit on the true peak stops the sound from clipping. - **Hardware encoder**: an encoder that runs on a graphics card or a video chip instead of the main processor. NVIDIA nvenc and Apple videotoolbox are the two that tidymedia supports. ## Where to next - `vignette("workflow")` shows a full research example. - `vignette("batch")` shows how to run a task function over many files. - `vignette("metadata")` shows how to read metadata into tibbles. - `vignette("verification")` shows how to check outputs and limit run time.