Content not found. Please use links in the navbar. # Page not found (404) # Get started with artoo This guide walks the whole artoo round-trip once, on the bundled demo data. A **spec** plus your **data** go through [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), write to a **file**, and read back **identical** — that loop is artoo’s lossless guarantee. Every step below runs as-is; there is nothing to download. ## The round-trip at a glance ![](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdib3g9IjAgMCAxMDQwIDI1MCIgd2lkdGg9IjEwMCUiIHJvbGU9ImltZyIgYXJpYS1sYWJlbD0iVGhlIGFydG9vIGxvc3NsZXNzIHJvdW5kLXRyaXA6IHNwZWMgcGx1cyBkYXRhIGdvIHRocm91Z2ggYXBwbHlfc3BlYyAoc2NhZmZvbGQsIGNvZXJjZSwgb3JkZXIsIHNvcnQsIHN0YW1wKSwgdGhlbiB3cml0ZSB0byBhIGZpbGUsIHRoZW4gcmVhZCBiYWNrIHRvIGlkZW50aWNhbCBkYXRhOyBzZXRfdHlwZSBhbmQgY2hlY2tfc3BlYyBmaXggYW5kIGluc3BlY3QgdGhlIHNwZWMsIGFuZCB0aGUgd2hvbGUgbG9vcCBpcyBsb3NzbGVzcywgc28gd2hhdCB5b3UgcmVhZCBiYWNrIGVxdWFscyB3aGF0IHlvdSB3cm90ZS4iPjxkZWZzPjxtYXJrZXIgaWQ9InJ0LWFycm93IiBtYXJrZXJ3aWR0aD0iMTAiIG1hcmtlcmhlaWdodD0iMTAiIHJlZng9IjciIHJlZnk9IjMiIG9yaWVudD0iYXV0byIgbWFya2VydW5pdHM9InVzZXJTcGFjZU9uVXNlIj48cGF0aCBkPSJNMCwwIEw3LDMgTDAsNiBaIiBmaWxsPSIjOTRhM2I4IiAvPjwvbWFya2VyPjxtYXJrZXIgaWQ9InJ0LWFycm93LWJsdWUiIG1hcmtlcndpZHRoPSIxMSIgbWFya2VyaGVpZ2h0PSIxMSIgcmVmeD0iNy41IiByZWZ5PSIzLjIiIG9yaWVudD0iYXV0byIgbWFya2VydW5pdHM9InVzZXJTcGFjZU9uVXNlIj48cGF0aCBkPSJNMCwwIEw3LjUsMy4yIEwwLDYuNCBaIiBmaWxsPSIjM2I4MmY2IiAvPjwvbWFya2VyPjwvZGVmcz48cmVjdCB4PSIxIiB5PSIxIiB3aWR0aD0iMTAzOCIgaGVpZ2h0PSIyNDgiIGZpbGw9IiNmZmZmZmYiIHN0cm9rZT0iI2VmZWZlZiIgc3Ryb2tlLXdpZHRoPSIyIiAvPjx0ZXh0IHg9Ijk5IiB5PSIzMiIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwgLWFwcGxlLXN5c3RlbSwgJiMzOTtTZWdvZSBVSSYjMzk7LCBzYW5zLXNlcmlmIiBmb250LXNpemU9IjEwLjUiIGZpbGw9IiM2NDc0OGIiPmZpeCDCtyBpbnNwZWN0IHRoZSBzcGVjPC90ZXh0Pjx0ZXh0IHg9Ijk5IiB5PSI1MSIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZm9udC1mYW1pbHk9InVpLW1vbm9zcGFjZSwgU0ZNb25vLVJlZ3VsYXIsIE1lbmxvLCBtb25vc3BhY2UiIGZvbnQtc2l6ZT0iMTIiIGZpbGw9IiMxZTI5M2IiPnNldF90eXBlKCkgwrcgY2hlY2tfc3BlYygpPC90ZXh0PjxsaW5lIHgxPSI5OSIgeTE9IjYwIiB4Mj0iOTkiIHkyPSI4MyIgc3Ryb2tlPSIjOTRhM2I4IiBzdHJva2Utd2lkdGg9IjEuNiIgbWFya2VyLWVuZD0idXJsKCNydC1hcnJvdykiPjwvbGluZT48ZyBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS13aWR0aD0iMS44Ij48bGluZSB4MT0iMTc2IiB5MT0iMTEyIiB4Mj0iMjA3IiB5Mj0iMTEyIiBtYXJrZXItZW5kPSJ1cmwoI3J0LWFycm93KSI+PC9saW5lPjxsaW5lIHgxPSIzODIiIHkxPSIxMTIiIHgyPSI0MTMiIHkyPSIxMTIiIG1hcmtlci1lbmQ9InVybCgjcnQtYXJyb3cpIj48L2xpbmU+PGxpbmUgeDE9IjUzOCIgeTE9IjExMiIgeDI9IjU2OSIgeTI9IjExMiIgbWFya2VyLWVuZD0idXJsKCNydC1hcnJvdykiPjwvbGluZT48bGluZSB4MT0iNjY2IiB5MT0iMTEyIiB4Mj0iNjk3IiB5Mj0iMTEyIiBtYXJrZXItZW5kPSJ1cmwoI3J0LWFycm93KSI+PC9saW5lPjxsaW5lIHgxPSI4MjIiIHkxPSIxMTIiIHgyPSI4NTMiIHkyPSIxMTIiIG1hcmtlci1lbmQ9InVybCgjcnQtYXJyb3cpIj48L2xpbmU+PC9nPjxyZWN0IHg9IjI0IiB5PSI4NiIgd2lkdGg9IjE1MCIgaGVpZ2h0PSI1MiIgZmlsbD0iI2Y5ZjlmOSIgc3Ryb2tlPSIjZTVlN2ViIiBzdHJva2Utd2lkdGg9IjEuNSIgLz48dGV4dCB4PSI5OSIgeT0iMTEyIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBkb21pbmFudC1iYXNlbGluZT0iY2VudHJhbCIgZm9udC1mYW1pbHk9InVpLW1vbm9zcGFjZSwgU0ZNb25vLVJlZ3VsYXIsIE1lbmxvLCBtb25vc3BhY2UiIGZvbnQtc2l6ZT0iMTQiIGZpbGw9IiMxZTI5M2IiPnNwZWMgKyBkYXRhPC90ZXh0PjxyZWN0IHg9IjIxMCIgeT0iODYiIHdpZHRoPSIxNzAiIGhlaWdodD0iNTIiIGZpbGw9IiMzYjgyZjYiIHN0cm9rZT0iIzI1NjNlYiIgc3Ryb2tlLXdpZHRoPSIxLjUiIC8+PHRleHQgeD0iMjk1IiB5PSIxMTIiIHRleHQtYW5jaG9yPSJtaWRkbGUiIGRvbWluYW50LWJhc2VsaW5lPSJjZW50cmFsIiBmb250LWZhbWlseT0idWktbW9ub3NwYWNlLCBTRk1vbm8tUmVndWxhciwgTWVubG8sIG1vbm9zcGFjZSIgZm9udC1zaXplPSIxNC41IiBmb250LXdlaWdodD0iNjAwIiBmaWxsPSIjZmZmZmZmIj5hcHBseV9zcGVjKCk8L3RleHQ+PHRleHQgeD0iMjk1IiB5PSIxNjAiIHRleHQtYW5jaG9yPSJtaWRkbGUiIGZvbnQtZmFtaWx5PSJzeXN0ZW0tdWksIC1hcHBsZS1zeXN0ZW0sICYjMzk7U2Vnb2UgVUkmIzM5Oywgc2Fucy1zZXJpZiIgZm9udC1zaXplPSIxMSIgZmlsbD0iIzY0NzQ4YiI+c2NhZmZvbGQgwrcgY29lcmNlIMK3IG9yZGVyIMK3IHNvcnQgwrcgc3RhbXA8L3RleHQ+PHJlY3QgeD0iNDE2IiB5PSI4NiIgd2lkdGg9IjEyMCIgaGVpZ2h0PSI1MiIgZmlsbD0iI2Y5ZjlmOSIgc3Ryb2tlPSIjZTVlN2ViIiBzdHJva2Utd2lkdGg9IjEuNSIgLz48dGV4dCB4PSI0NzYiIHk9IjExMiIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZG9taW5hbnQtYmFzZWxpbmU9ImNlbnRyYWwiIGZvbnQtZmFtaWx5PSJ1aS1tb25vc3BhY2UsIFNGTW9uby1SZWd1bGFyLCBNZW5sbywgbW9ub3NwYWNlIiBmb250LXNpemU9IjE0IiBmaWxsPSIjMWUyOTNiIj53cml0ZV8qKCk8L3RleHQ+PHJlY3QgeD0iNTcyIiB5PSI4NiIgd2lkdGg9IjkyIiBoZWlnaHQ9IjUyIiBmaWxsPSIjZWVmM2ZiIiBzdHJva2U9IiNjN2RiZmYiIHN0cm9rZS13aWR0aD0iMS41IiAvPjx0ZXh0IHg9IjYxOCIgeT0iMTEyIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBkb21pbmFudC1iYXNlbGluZT0iY2VudHJhbCIgZm9udC1mYW1pbHk9InVpLW1vbm9zcGFjZSwgU0ZNb25vLVJlZ3VsYXIsIE1lbmxvLCBtb25vc3BhY2UiIGZvbnQtc2l6ZT0iMTQiIGZpbGw9IiMxZTNhNWYiPmZpbGU8L3RleHQ+PHJlY3QgeD0iNzAwIiB5PSI4NiIgd2lkdGg9IjEyMCIgaGVpZ2h0PSI1MiIgZmlsbD0iI2Y5ZjlmOSIgc3Ryb2tlPSIjZTVlN2ViIiBzdHJva2Utd2lkdGg9IjEuNSIgLz48dGV4dCB4PSI3NjAiIHk9IjExMiIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZG9taW5hbnQtYmFzZWxpbmU9ImNlbnRyYWwiIGZvbnQtZmFtaWx5PSJ1aS1tb25vc3BhY2UsIFNGTW9uby1SZWd1bGFyLCBNZW5sbywgbW9ub3NwYWNlIiBmb250LXNpemU9IjE0IiBmaWxsPSIjMWUyOTNiIj5yZWFkXyooKTwvdGV4dD48cmVjdCB4PSI4NTYiIHk9Ijg2IiB3aWR0aD0iMTYwIiBoZWlnaHQ9IjUyIiBmaWxsPSIjZjlmOWY5IiBzdHJva2U9IiNlNWU3ZWIiIHN0cm9rZS13aWR0aD0iMS41IiAvPjx0ZXh0IHg9IjkzNiIgeT0iMTEyIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBkb21pbmFudC1iYXNlbGluZT0iY2VudHJhbCIgZm9udC1mYW1pbHk9InVpLW1vbm9zcGFjZSwgU0ZNb25vLVJlZ3VsYXIsIE1lbmxvLCBtb25vc3BhY2UiIGZvbnQtc2l6ZT0iMTQiIGZpbGw9IiMxZTI5M2IiPmlkZW50aWNhbCBkYXRhPC90ZXh0PjxwYXRoIGQ9Ik05MzYsMTQwIFYyMDIgSDk5IFYxNDMiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzNiODJmNiIgc3Ryb2tlLXdpZHRoPSIxLjgiIHN0cm9rZS1kYXNoYXJyYXk9IjYgNSIgbWFya2VyLWVuZD0idXJsKCNydC1hcnJvdy1ibHVlKSIgLz48dGV4dCB4PSI1MTciIHk9IjIyOCIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwgLWFwcGxlLXN5c3RlbSwgJiMzOTtTZWdvZSBVSSYjMzk7LCBzYW5zLXNlcmlmIiBmb250LXNpemU9IjEzIiBmb250LXdlaWdodD0iNjAwIiBmaWxsPSIjM2I4MmY2Ij5sb3NzbGVzcyByb3VuZC10cmlwIOKAlCB3aGF0IHlvdSByZWFkIGJhY2sgZXF1YWxzIHdoYXQgeW91IHdyb3RlPC90ZXh0Pjwvc3ZnPg==) ## 1. Get a spec A `artoo_spec` is the canonical description of your datasets: variables, CDISC data types, lengths, labels, controlled-terminology codelists, and sort keys — always for exactly **one** CDISC standard. [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) reads one from Define-XML, a Pinnacle 21 workbook, or artoo’s native JSON; [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) assembles one from metadata frames. The package bundles ready-made specs built from the official CDISC Define-XML 2.1 release examples — `adam_spec` for ADaM (ADSL, ADAE) and `sdtm_spec` for SDTM (DM, VS, TS, SUPPDM). Each also ships as a P21 workbook you can open in Excel: ``` r adam_spec ``` Study: CDISC-Sample Standard: ADaMIG 1.1 Datasets: 2 Variables: 104 Codelists: 30 Methods: 54 Comments: 22 Documents: 9 Spec for: ADSL, ADAE ``` r p21 <- system.file("extdata", "adam-spec.xlsx", package = "artoo") identical(spec_standard(read_spec(p21)), spec_standard(adam_spec)) ``` [1] TRUE Because [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) and [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) are inverses on each format, format conversion is one composition — Define-XML in, P21 workbook out: ``` r read_spec("define.xml") |> write_spec("spec.xlsx") ``` ## 2. Apply the spec [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) is the conform pipeline: it coerces each column to its CDISC data type, orders the columns, sorts by the dataset keys, and stamps the result with its metadata. A variable the spec declares but the data lacks is reported, never fabricated as an empty column. The input is never mutated, no column is ever dropped, and a coercion that would damage values aborts before it runs — with two honest one-line exits: keep the wider source type with `apply_spec(..., on_coercion_loss = "keep")`, or retype the spec with [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md). ``` r adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL") ``` 6 variables the spec declares are absent from the data (not added): `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. ℹ See `conformance(x)` for the findings. The conformance findings ride along on the result — read them back as a frame with [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md): ``` r nrow(conformance(adsl)) ``` [1] 12 The pipeline is standard-neutral: an SDTM domain conforms identically — only the spec and the dataset change. ``` r dm <- apply_spec(cdisc_dm, sdtm_spec, "DM") ``` 1 variable the spec declares is absent from the data (not added): `BRTHDTC`. ℹ See `conformance(x)` for the findings. ``` r nrow(conformance(dm)) ``` [1] 14 ## 3. Inspect the columns [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) is the quick look a SAS programmer expects from `PROC CONTENTS`: one row per variable with position, type, length, format, label, and the CDISC key sequence. It works on a conformed frame, any plain data frame, or a file path: ``` r columns(adsl) ``` ADSL -- 48 variables, 60 obs # Variable Type Len Format Label Key 1 STUDYID Char 12 Study Identifier 1 2 USUBJID Char 11 Unique Subject Identifier 2 3 SUBJID Char 4 Subject Identifier for the Study 4 SITEID Char 3 Study Site Identifier 5 SITEGR1 Char 3 Pooled Site Group 1 6 ARM Char 20 Description of Planned Arm 7 TRT01P Char 20 Planned Treatment for Period 01 8 TRT01PN Num Planned Treatment for Period 01 (N) 9 TRT01A Char 20 Actual Treatment for Period 01 10 TRT01AN Num Actual Treatment for Period 01 (N) 11 TRTSDT Num DATE9. Date of First Exposure to Treatment 12 TRTEDT Num DATE9. Date of Last Exposure to Treatment 13 AVGDD Num 5.1 Avg Daily Dose (as planned) 14 CUMDOSE Num 8.1 Cumulative Dose (as planned) 15 AGE Num Age 16 AGEGR1 Char 5 Pooled Age Group 1 17 AGEGR1N Num Pooled Age Group 1 (N) 18 AGEU Char 5 Age Units 19 RACE Char 32 Race 20 RACEN Num Race (N) 21 SEX Char 1 Sex 22 ETHNIC Char 22 Ethnicity 23 SAFFL Char 1 Safety Population Flag 24 ITTFL Char 1 Intent-To-Treat Population Flag 25 EFFFL Char 1 Efficacy Population Flag 26 COMP8FL Char 1 Completers of Week 8 Population Flag 27 COMP16FL Char 1 Completers of Week 16 Population Flag 28 COMP24FL Char 1 Completers of Week 24 Population Flag 29 DISCONFL Char 1 Subject Discontinued Study Flag 30 DSRAEFL Char 1 Subject Discontinued due to AE Flag 31 DTHFL Char 1 Subject Death Flag 32 BMIBL Num 5.1 Baseline BMI (kg/m^2) 33 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 34 HEIGHTBL Num 6.1 Baseline Height (cm) 35 WEIGHTBL Num 6.1 Baseline Weight (kg) 36 EDUCLVL Num Years of Education 37 DURDIS Num 6.1 Duration of Disease (Months) 38 DURDSGR1 Char 4 Pooled Disease Duration Group 1 39 VISIT1DT Num DATE9. Date of Visit 1 40 RFSTDTC Char 10 Subject Reference Start Date/Time 41 RFENDTC Char 10 Subject Reference End Date/Time 42 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) 43 RFENDT Num Date of Discontinuation/Completion 44 TRTDUR Num 45 DISONSDT Num DATE9. 46 DCDECOD Char 27 47 DCREASCD Char 18 48 MMSETOT Num ## 4. Write to any format — losslessly Every writer carries the full metadata model, so the write is lossless by construction. The writers return their input invisibly, so one conformed frame fans out to every deliverable: ``` r xpt <- tempfile(fileext = ".xpt") json <- tempfile(fileext = ".json") adsl |> write_xpt(xpt) |> write_json(json) ``` Any file converts to any other without re-applying the spec — the metadata travels inside (or beside) the container: ``` r parquet <- tempfile(fileext = ".parquet") write_parquet(read_json(json), parquet) ``` ## 5. Read back, intact Reading restores the values, the R classes (dates as `Date`, times as `hms`), the labels, and the metadata — identically from every format: ``` r back <- read_json(json) get_meta(back)@dataset$records ``` [1] 60 ``` r columns(back) ``` ADSL -- 48 variables, 60 obs # Variable Type Len Format Label Key 1 STUDYID Char 12 Study Identifier 1 2 USUBJID Char 11 Unique Subject Identifier 2 3 SUBJID Char 4 Subject Identifier for the Study 4 SITEID Char 3 Study Site Identifier 5 SITEGR1 Char 3 Pooled Site Group 1 6 ARM Char 20 Description of Planned Arm 7 TRT01P Char 20 Planned Treatment for Period 01 8 TRT01PN Num Planned Treatment for Period 01 (N) 9 TRT01A Char 20 Actual Treatment for Period 01 10 TRT01AN Num Actual Treatment for Period 01 (N) 11 TRTSDT Num DATE9. Date of First Exposure to Treatment 12 TRTEDT Num DATE9. Date of Last Exposure to Treatment 13 AVGDD Num 5.1 Avg Daily Dose (as planned) 14 CUMDOSE Num 8.1 Cumulative Dose (as planned) 15 AGE Num Age 16 AGEGR1 Char 5 Pooled Age Group 1 17 AGEGR1N Num Pooled Age Group 1 (N) 18 AGEU Char 5 Age Units 19 RACE Char 32 Race 20 RACEN Num Race (N) 21 SEX Char 1 Sex 22 ETHNIC Char 22 Ethnicity 23 SAFFL Char 1 Safety Population Flag 24 ITTFL Char 1 Intent-To-Treat Population Flag 25 EFFFL Char 1 Efficacy Population Flag 26 COMP8FL Char 1 Completers of Week 8 Population Flag 27 COMP16FL Char 1 Completers of Week 16 Population Flag 28 COMP24FL Char 1 Completers of Week 24 Population Flag 29 DISCONFL Char 1 Subject Discontinued Study Flag 30 DSRAEFL Char 1 Subject Discontinued due to AE Flag 31 DTHFL Char 1 Subject Death Flag 32 BMIBL Num 5.1 Baseline BMI (kg/m^2) 33 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 34 HEIGHTBL Num 6.1 Baseline Height (cm) 35 WEIGHTBL Num 6.1 Baseline Weight (kg) 36 EDUCLVL Num Years of Education 37 DURDIS Num 6.1 Duration of Disease (Months) 38 DURDSGR1 Char 4 Pooled Disease Duration Group 1 39 VISIT1DT Num DATE9. Date of Visit 1 40 RFSTDTC Char 10 Subject Reference Start Date/Time 41 RFENDTC Char 10 Subject Reference End Date/Time 42 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) 43 RFENDT Num Date of Discontinuation/Completion 44 TRTDUR Num 45 DISONSDT Num DATE9. 46 DCDECOD Char 27 47 DCREASCD Char 18 48 MMSETOT Num One honest caveat: the XPORT byte layout stores only name, label, length, and formats, so [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) on an `.xpt` path shows a blank `Key` — the key sequence (like codelist references) rides the metadata-carrying formats and the in-session frame, never the 1980s transport bytes. That round-trip identity is the whole point: what you submit is what you archived is what you analysed. ## Where to next - [Specifications](https://vthanik.github.io/artoo/articles/specs.html) — read a spec from Define-XML or a workbook, inspect it with the `spec_*` accessors, and fix it in place with [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) / [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md). - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.html) — [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) in depth, then every conformance finding from [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) and [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md), and the errors artoo raises. - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.html) — any-to-any round trips, encodings, the `on_invalid` policy, and the qualification evidence a regulated pipeline needs. - [Recipes](https://vthanik.github.io/artoo/articles/recipes.html) — end-to-end ADaM and SDTM builds, dates and `--DTC`, and codelist decoding, each rendered live on the demo data. \`\`\` # Conform & validate [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) is the heart of artoo: it conforms a raw frame to a spec — coerce, order, sort, stamp — and never silently damages data. Validation is the same surface read the other way: [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) and [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) return every finding at once, and every abort artoo raises is a classed condition that names its fix. This article covers both. ## 1. Conform with `apply_spec()` [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) runs the same five steps, in the same documented order — scaffold the spec’s columns, coerce each to its CDISC data type, order to the spec, sort by the dataset keys, stamp the metadata — and returns a frame ready to write. A variable the spec declares but the data lacks is reported, never fabricated as an empty column: ``` r adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL") ``` 6 variables the spec declares are absent from the data (not added): `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. ℹ See `conformance(x)` for the findings. Two arguments cover the common cases. `extra = "drop"` trims columns the spec does not mention — announced, and recorded by the `extra_variable` finding, so the drop is never silent: ``` r adsl_lean <- apply_spec(cdisc_adsl, adam_spec, "ADSL", extra = "drop") ``` 6 variables the spec declares are absent from the data (not added): `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. ℹ See `conformance(x)` for the findings. Dropped 5 undeclared variables: `TRTDUR`, `DISONSDT`, `DCDECOD`, `DCREASCD`, and `MMSETOT` ``` r ncol(adsl_lean) <= ncol(cdisc_adsl) ``` [1] TRUE `na_position` controls where missing key values sort. The default `"first"` matches SAS `PROC SORT` (and FDA submission datasets); set `"last"` only when your comparison target is R’s [`order()`](https://rdrr.io/r/base/order.html): ``` r nrow(apply_spec(cdisc_adsl, adam_spec, "ADSL", na_position = "last")) ``` 6 variables the spec declares are absent from the data (not added): `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. ℹ See `conformance(x)` for the findings. [1] 60 ## 2. Lossless or loud artoo’s core rule is that no operation silently damages data. A coercion that would truncate fractions or overflow R’s 32-bit integer range aborts with `artoo_error_type` **before** touching a value — and this gate is independent of `conformance`, so turning conformance off does not bypass it: ``` r vars <- spec_variables(adam_spec) vars$data_type[vars$variable == "AGE"] <- "integer" strict <- artoo_spec( adam_spec@datasets, vars, codelists = adam_spec@codelists, study = spec_study(adam_spec) ) raw <- cdisc_adsl raw$AGE[1] <- raw$AGE[1] + 0.5 apply_spec(raw, strict, "ADSL", conformance = "off") ``` Error: ! Coercion to the spec dataTypes would lose data. ✖ Integer coercion would truncate fractional values in: AGE (1). ℹ This gate is separate from `conformance`; `conformance = "off"` does not bypass it. ℹ To keep these values in R, set `apply_spec(on_coercion_loss = "keep")`, or retype the spec with `set_type()` (dataType "float" or "decimal"). ℹ To see every finding at once, run `check_spec(x, spec, dataset)`. You have two honest one-line exits: keep the wider source type with `apply_spec(on_coercion_loss = "keep")` (the value is preserved and the mismatch is left as an `integer_fraction` finding), or retype the spec with [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md). ## 3. See every finding, without halting [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) stops at the first thing that would lose data; to *list* every issue instead, ask. [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) returns all of them as a tidy frame — one row per finding, no abort: ``` r findings <- check_spec(cdisc_adsl, adam_spec, "ADSL") findings ``` ADSL: 0 errors, 11 warnings, 15 notes Warnings -------- [missing_permissible] Permissible spec variable 'TRTDURD' is absent from the data. (ADSL.TRTDURD) [missing_permissible] Permissible spec variable 'DISONDT' is absent from the data. (ADSL.DISONDT) [missing_permissible] Permissible spec variable 'EOSSTT' is absent from the data. (ADSL.EOSSTT) [missing_permissible] Permissible spec variable 'DCSREAS' is absent from the data. (ADSL.DCSREAS) [missing_permissible] Permissible spec variable 'EOSDISP' is absent from the data. (ADSL.EOSDISP) [missing_permissible] Permissible spec variable 'MMS1TSBL' is absent from the data. (ADSL.MMS1TSBL) [extra_variable] Column 'TRTDUR' is not declared in the spec. (ADSL.TRTDUR) [extra_variable] Column 'DISONSDT' is not declared in the spec. (ADSL.DISONSDT) [extra_variable] Column 'DCDECOD' is not declared in the spec. (ADSL.DCDECOD) [extra_variable] Column 'DCREASCD' is not declared in the spec. (ADSL.DCREASCD) [extra_variable] Column 'MMSETOT' is not declared in the spec. (ADSL.MMSETOT) Notes ----- [type_mismatch] 'TRT01PN' is stored as double but the spec dataType 'integer' wants integer. (ADSL.TRT01PN) [type_mismatch] 'TRT01AN' is stored as double but the spec dataType 'integer' wants integer. (ADSL.TRT01AN) [type_mismatch] 'TRTSDT' is stored as double but the spec dataType 'integer' wants integer. (ADSL.TRTSDT) [type_mismatch] 'TRTEDT' is stored as double but the spec dataType 'integer' wants integer. (ADSL.TRTEDT) [type_mismatch] 'AGE' is stored as double but the spec dataType 'integer' wants integer. (ADSL.AGE) [type_mismatch] 'AGEGR1N' is stored as double but the spec dataType 'integer' wants integer. (ADSL.AGEGR1N) [type_mismatch] 'RACEN' is stored as double but the spec dataType 'integer' wants integer. (ADSL.RACEN) [label_match] 'DISCONFL' label 'Did the Subject Discontinue the Study?' differs from the spec label 'Subject Discontinued Study Flag'. (ADSL.DISCONFL) [label_match] 'DSRAEFL' label 'Discontinued due to AE?' differs from the spec label 'Subject Discontinued due to AE Flag'. (ADSL.DSRAEFL) [label_match] 'DTHFL' label 'Subject Died?' differs from the spec label 'Subject Death Flag'. (ADSL.DTHFL) [codelist_membership_extensible] 'BMIBLGR1' has 1 value(s) outside extensible codelist 'CL.BMICAT': >=30. (ADSL.BMIBLGR1) [type_mismatch] 'EDUCLVL' is stored as double but the spec dataType 'integer' wants integer. (ADSL.EDUCLVL) [type_mismatch] 'VISIT1DT' is stored as double but the spec dataType 'integer' wants integer. (ADSL.VISIT1DT) [type_mismatch] 'VISNUMEN' is stored as double but the spec dataType 'integer' wants integer. (ADSL.VISNUMEN) [type_mismatch] 'RFENDT' is stored as double but the spec dataType 'integer' wants integer. (ADSL.RFENDT) The same report rides along on a conformed result under the default `conformance = "warn"`; read it back with [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md): ``` r nrow(conformance(adsl)) ``` [1] 12 Scaling to a whole study is one call: [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) takes a named list of datasets and returns the same shape, so “is my study submittable?” has a single answer. ``` r nrow(check_study(adam_spec, list(ADSL = cdisc_adsl))) ``` [1] 26 ## 4. Scope the checks [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) toggles each conformance dimension on or off; pass the result to [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) / [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) to narrow a run to what you care about (here, everything except the type-mismatch note): ``` r ck <- artoo_checks(type_mismatch = FALSE) nrow(check_spec(cdisc_adsl, adam_spec, "ADSL", checks = ck)) ``` [1] 15 ## 5. The error family Every abort artoo raises carries a class of the form `artoo_error_` (plus `artoo_error` and `artoo_condition`, so one `tryCatch(artoo_error = )` catches them all), and the data-protection conditions carry their evidence as data, so a qualification harness can assert on `cnd$variables` / `cnd$findings` rather than match message text. The kinds, each triggered live: `artoo_error_input` — the call is malformed (here, an unknown dataset): ``` r apply_spec(cdisc_dm, sdtm_spec, "NOPE") ``` Error: ! `dataset` must be one of the spec's datasets. ✖ "NOPE" is not in the spec. ℹ Available: "TS", "DM", "VS", and "SUPPDM". `artoo_error_conformance` — `apply_spec(conformance = "abort")` met error-severity findings: ``` r dm <- cdisc_dm dm$SEX[1] <- "X9" apply_spec(dm, sdtm_spec, "DM", conformance = "abort") ``` Error: ! Data does not conform to the spec for "DM". ✖ 'SEX' has 1 value(s) outside codelist 'CL.SEX': X9. `artoo_error_codelist` — [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md) met a value outside the codelist’s terms (pass `no_match = "keep"` / `"na"` to carry it through): ``` r decode_column(dm, sdtm_spec, "DM", from = "SEX", to = "SEXDECD") ``` Error: ! Values in `SEX` are not in codelist "CL.SEX". ✖ Unmatched: "X9". ℹ Set `no_match = "keep"` or `"na"` to allow them. `artoo_error_codec` — the bytes cannot travel (an invalid-UTF-8 value for Dataset-JSON); re-read with the true `encoding=`, or set an `on_invalid` policy on the writer: ``` r clean <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") clean$USUBJID[1] <- rawToChar(as.raw(c(0x63, 0xE9))) write_json(clean, tempfile(fileext = ".json")) ``` Error in `write_json()`: ! Cannot encode 1 value as UTF-8. ✖ Invalid bytes (hex-escaped): "c". ℹ Re-read the source with the correct `encoding`, or set `on_invalid`. (The remaining kind, `artoo_error_spec`, fires at spec construction when a slot references an unknown dataset or codelist, mixes standards, or duplicates a definition.) ## Where to next - [Specifications](https://vthanik.github.io/artoo/articles/specs.md) — build, inspect, and repair the spec this verb consumes. - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.md) — write the conformed frame to any format, and the qualification evidence behind “lossless”. - [Recipes](https://vthanik.github.io/artoo/articles/recipes.md) — conform inside an end-to-end ADaM and SDTM build. - [Get started](https://vthanik.github.io/artoo/articles/artoo.md) — the round-trip from the top. # Formats & lossless conversion artoo’s conversion story is one sentence: every codec reads and writes the same canonical metadata, so any format converts to any other without loss. This article shows the round trips, what “lossless” means concretely, how the encoding policy keeps a single bad byte from deciding which formats your dataset can travel in, and the evidence a regulated pipeline can point to. ## 1. One metadata model, four carriers A conformed frame carries its `artoo_meta`; each writer embeds it in the format’s own idiom (XPORT NAMESTR records, the Dataset-JSON itemGroup, a Parquet key-value sidecar, rds attributes): ``` r dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") ``` 1 variable the spec declares is absent from the data (not added): `BRTHDTC`. ``` r xpt <- tempfile(fileext = ".xpt") json <- tempfile(fileext = ".json") parquet <- tempfile(fileext = ".parquet") dm |> write_xpt(xpt) |> write_json(json) |> write_parquet(parquet) ``` Warning in write_xpt(dm, xpt): Widened 1 column past the declared spec length: "STUDYID (7 -> 12)". ℹ Values need more bytes than the spec length; data was kept whole. ℹ Update the spec length, or shorten the data, so the file matches its declared metadata. Conversion is read one, write the other — and the metadata survives the hop: ``` r from_xpt <- read_xpt(xpt) identical(get_meta(from_xpt)@columns, get_meta(read_json(json))@columns) ``` [1] FALSE One caveat is structural, not artoo’s: XPORT v5 cannot carry `keySequence`, codelist references, or origin (its NAMESTR record has no field for them), so an `.xpt`-sourced frame shows a blank Key pane by design. Route through Dataset-JSON or Parquet when those must survive. ## 2. Encodings: recorded, inherited, never silently transformed artoo text is always UTF-8 in memory; encodings matter only at the file boundary. Each reader records the source encoding in the metadata, and `write_*(encoding = NULL)` inherits it, so a round trip is byte-faithful. The name rosetta lives in [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md): ``` r enc <- artoo_encodings() enc[enc$sas == "WLATIN1", ] ``` r sas python 3 windows-1252 WLATIN1 cp1252 description 3 Western European Windows; the usual US/EU SAS session (WLATIN1) ### Reading a file whose bytes are not UTF-8 `encoding=` on a reader is a *transcode*, not a label. Say a Parquet file was written by a system whose text is Windows-1252 (SAS `WLATIN1`) rather than the UTF-8 the Parquet spec assumes: ``` r latin1 <- function(x) { out <- iconv(x, "UTF-8", "windows-1252") Encoding(out) <- "unknown" out } foreign <- data.frame( SITEID = latin1(c("PARIS", "MÜNCHEN")), INVNAM = latin1(c("Curie", "Öztürk")) ) pq <- tempfile(fileext = ".parquet") nanoparquet::write_parquet(foreign, pq) ``` Naming the source charset converts the bytes to UTF-8 and normalises them to NFC, so what lands in R is ordinary text — printing, [`View()`](https://rdrr.io/r/utils/View.html), [`nchar()`](https://rdrr.io/r/base/nchar.html), regex, and collation all behave, and every writer downstream re-encodes from there: ``` r site <- read_parquet(pq, encoding = "wlatin1") site ``` SITEID INVNAM 1 PARIS Curie 2 MÜNCHEN Öztürk ``` r nchar(site$SITEID) # characters, not bytes ``` [1] 5 7 ``` r grepl("Ü", site$SITEID) ``` [1] FALSE TRUE Two ways it can look wrong, both diagnostic rather than silent: *Mojibake* (`MÜNCHEN`) means you declared an `encoding=` for a file that was already UTF-8 — drop the argument. A hex escape (`caf`) means a byte that is not valid in the charset you named; artoo escapes it rather than dropping it, so the corruption stays visible. And omitting `encoding=` on a genuinely non-UTF-8 file is not a silent mis-read either — the reader refuses it: ``` r read_parquet(pq) ``` Error in `read_parquet()`: ! Could not read '/tmp/RtmpBzuEaV/file1e6527f95aa5.parquet' as "parquet". ✖ entry 2 has wrong Encoding; marked as "UTF-8" but leading byte 0xDC followed by invalid continuation byte (0x4E) at position 3 ## 3. `on_invalid`: one policy, all four writers A value that cannot be represented in the target — an unencodable character for xpt’s target charset, or bytes that are not valid UTF-8 for Dataset-JSON / NDJSON / Parquet — hits the same policy everywhere: `"error"` (default), `"replace"`, `"ignore"`, or `"translit"` (fold smart punctuation to its exact ASCII form; see the [WLATIN1 → UTF-8 migration article](https://vthanik.github.io/artoo/articles/migrate-encoding.md)). The default names the offenders hex-escaped: ``` r bad <- dm bad$USUBJID[1] <- rawToChar(as.raw(c(0x63, 0xE9))) # a stray latin1 byte write_json(bad, tempfile(fileext = ".json")) ``` Error in `write_json()`: ! Cannot encode 1 value as UTF-8. ✖ Invalid bytes (hex-escaped): "c". ℹ Re-read the source with the correct `encoding`, or set `on_invalid`. `"replace"` substitutes `?` and warns, so the write completes and the file is valid: ``` r fixed <- tempfile(fileext = ".json") write_json(bad, fixed, on_invalid = "replace") ``` Warning in write_json(bad, fixed, on_invalid = "replace"): Replaced invalid UTF-8 bytes with "?" in 1 value. ℹ Use `on_invalid = "error"` to fail loudly instead. ``` r read_json(fixed)$USUBJID[1] ``` [1] "c?" The right fix is upstream — re-read the source with its true `encoding=` so the bytes arrive as the characters they were — but the policy means a mis-declared byte in one record no longer makes a dataset “submittable as XPT but not as Dataset-JSON”, or vice versa. ## 4. Qualification: the evidence behind “lossless” The first question a pharma organization asks of any package in a submission pipeline is “can we qualify this?”. None of the below is a regulatory claim — qualification is your process — but every property is machine-checkable in your own environment. **Lossless or loud.** Read-write-read returns the same values and the same metadata; a lossy coercion or an unencodable byte aborts rather than damaging data; even the explicit `extra = "drop"` trim announces itself and leaves a finding. The data-protection conditions carry their evidence as data, so a harness asserts on it programmatically: ``` r vars <- spec_variables(adam_spec) vars$data_type[vars$variable == "AGE"] <- "integer" strict <- artoo_spec( adam_spec@datasets, vars, codelists = adam_spec@codelists, study = spec_study(adam_spec) ) raw <- cdisc_adsl raw$AGE[1] <- raw$AGE[1] + 0.5 tryCatch( apply_spec(raw, strict, "ADSL", conformance = "off"), artoo_error_type = function(cnd) cnd$variables ) ``` variable data_type n reason 1 AGE integer 1 truncated **The development gates**, enforced on every change and in CI: `R CMD check` at 0 errors / 0 warnings / 0 notes; a test suite in the thousands of assertions, including byte-level golden files for the codecs, a cross-format round-trip matrix, and fuzzed-input tests that assert every failure is a classed artoo condition; line coverage of at least 95% on every file in `R/`; and reproducible demo data, every bundled object rebuilt by script from public, checksum-pinned sources. **Standards the behavior is pinned to**, cited in the docs rather than re-invented: CDISC Dataset-JSON v1.1 (type vocabulary, UTF-8), the FDA Study Data Technical Conformance Guide (XPORT expectations, ASCII gate), SAS XPORT v5/v8 (TS-140/TS-340), RFC 8259, IANA character set names, and Unicode NFC (UAX \#15). ## Where to next - [Specifications](https://vthanik.github.io/artoo/articles/specs.md) — the spec whose metadata every codec carries. - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.md) — produce the conformed frame these writers persist. - [Recipes](https://vthanik.github.io/artoo/articles/recipes.md) — conversion inside an end-to-end ADaM and SDTM build. - [Get started](https://vthanik.github.io/artoo/articles/artoo.md) — the round-trip from the top. # Articles ### Articles - [Specifications](https://vthanik.github.io/artoo/articles/specs.md): - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.md): - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.md): - [Migrating clinical data from WLATIN1 to UTF-8](https://vthanik.github.io/artoo/articles/migrate-encoding.md): - [Recipes](https://vthanik.github.io/artoo/articles/recipes.md): # Migrating clinical data from WLATIN1 to UTF-8 Most legacy clinical data was produced by SAS sessions running WLATIN1 (Windows-1252); modern pipelines — Dataset-JSON, Parquet, R itself — are UTF-8. The two agree on every ASCII character and disagree everywhere else, so a migration has exactly three failure modes: mojibake (bytes interpreted as the wrong charset), truncation (byte-counted lengths that multibyte UTF-8 overflows), and unrepresentable characters (an ASCII target that cannot carry what UTF-8 can). This article shows how artoo handles each one. ## 1. Reading WLATIN1 data [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) detects the encoding: if every byte in the file validates as UTF-8 it reads as UTF-8, otherwise it assumes Windows-1252 — the right default for legacy clinical files. Formats whose spec pins UTF-8 (Dataset-JSON, Parquet) are read as UTF-8 and *refuse loudly* when the bytes disagree, rather than guessing: ``` r latin1 <- function(x) { out <- iconv(x, "UTF-8", "windows-1252") Encoding(out) <- "unknown" out } foreign <- data.frame(SITEID = latin1(c("PARIS", "MÜNCHEN"))) pq <- tempfile(fileext = ".parquet") nanoparquet::write_parquet(foreign, pq) read_parquet(pq) ``` Error in `read_parquet()`: ! Could not read '/tmp/RtmpNJ7pbJ/file1e8257749ccc.parquet' as "parquet". ✖ entry 2 has wrong Encoding; marked as "UTF-8" but leading byte 0xDC followed by invalid continuation byte (0x4E) at position 3 That refusal is the prompt to name the true source charset. `encoding=` is a *transcode*, not a label — the bytes are converted to UTF-8 and normalised (NFC), so what lands in R is ordinary text: ``` r site <- read_parquet(pq, encoding = "wlatin1") site$SITEID ``` [1] "PARIS" "MÜNCHEN" Any SAS or IANA spelling works — `"wlatin1"`, `"cp1252"`, `"windows-1252"` are the same charset; so are the OEM/DOS names (`"pcoem850"`) a very old archive might declare. The rosetta is [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md). Note that `"latin1"` (ISO-8859-1) is **not** WLATIN1: the two differ across the 0x80–0x9F range that holds the Euro sign and every smart quote. ## 2. Catching a mis-declared read The classic silent corruption is reading WLATIN1 bytes *as if* they were UTF-8 (in SAS terms: a data set whose encoding attribute lies). Those bytes do not validate as UTF-8, and the `invalid_encoding` conformance dimension flags them before any writer runs: ``` r bad <- cdisc_dm bad$ETHNIC <- rawToChar(as.raw(c(0x4D, 0xDC, 0x4E))) # WLATIN1 bytes, raw f <- check_spec(bad, sdtm_spec, "DM") f[f$check == "invalid_encoding", c("variable", "severity", "message")] ``` variable severity 14 ETHNIC error message 14 'ETHNIC' has 60 value(s) whose bytes are not valid UTF-8; re-read the source with its true encoding. The fix is always upstream — re-read the source with its true `encoding=` — never a re-label. ## 3. Lengths are bytes, and UTF-8 needs more of them A WLATIN1 character is always one byte; in UTF-8 the same character can take up to four. A spec length calibrated on WLATIN1 data (`Ö` = 1 byte) under-declares the UTF-8 form (`Ö` = 2 bytes). SAS handles this with the CVP engine and a guessed multiplier; artoo measures the actual bytes, never truncates, and tells you when the declared length had to give: ``` r sp <- artoo_spec( data.frame(dataset = "DM"), data.frame( dataset = "DM", variable = "INVNAM", data_type = "string", length = 6L ) ) x <- apply_spec( data.frame(INVNAM = "Öztürk"), sp, "DM", conformance = "off" ) xpt <- tempfile(fileext = ".xpt") write_xpt(x, xpt) ``` Warning in write_xpt(x, xpt): Widened 1 column past the declared spec length: "INVNAM (6 -> 8)". ℹ Values need more bytes than the spec length; data was kept whole. ℹ Update the spec length, or shorten the data, so the file matches its declared metadata. The `length_overflow` check reports the same fact at conformance time, so a migration can fix the spec (the durable answer) instead of relying on the writer’s defence. ## 4. Smart punctuation: the WLATIN1 characters ASCII can fold Text that passed through a word processor carries typographic punctuation — curly quotes, en/em dashes, ellipses. In WLATIN1 these live in rows 8–9 of the code page (one byte each); in UTF-8 they are three bytes each; in US-ASCII they do not exist at all. They are also the one class of character with an *exact* ASCII equivalent, published by SAS as the NLS punctuation fold: | WLATIN1 (hex) | Unicode | Character | Description | ASCII fold | |---------------|---------|:---------:|-----------------------------------|:----------:| | `82` | U+201A | ‚ | single low-9 quotation mark | `,` | | `84` | U+201E | „ | double low-9 quotation mark | `"` | | `85` | U+2026 | … | horizontal ellipsis | `...` | | `8B` | U+2039 | ‹ | single left-pointing angle quote | `<` | | `91` | U+2018 | ‘ | left single quotation mark | `'` | | `92` | U+2019 | ’ | right single quotation mark | `'` | | `93` | U+201C | “ | left double quotation mark | `"` | | `94` | U+201D | ” | right double quotation mark | `"` | | `95` | U+2022 | • | bullet | `*` | | `96` | U+2013 | – | en dash | `-` | | `97` | U+2014 | — | em dash | `-` | | `9B` | U+203A | › | single right-pointing angle quote | `>` | Every writer accepts `on_invalid = "translit"`, which applies exactly this table and nothing else: ``` r note <- data.frame( USUBJID = "01-701-1015", COMMENT = "Patient’s dose – “held”…" ) ascii_xpt <- tempfile(fileext = ".xpt") write_xpt(note, ascii_xpt, encoding = "US-ASCII", on_invalid = "translit") ``` Warning in write_xpt(note, ascii_xpt, encoding = "US-ASCII", on_invalid = "translit"): Transliterated smart punctuation to ASCII in 1 value for "US-ASCII". ℹ The fold follows the SAS NLS punctuation table (quotes, dashes, ellipsis, bullet). ``` r read_xpt(ascii_xpt)$COMMENT ``` [1] "Patient's dose - \"held\"..." `"translit"` is deliberately narrower than `"replace"`: punctuation folds because the ASCII form carries the same meaning; a character with no equivalent — a diacritic in a name — still aborts, because `"Öztürk"` silently becoming `"Ozturk"` (or worse, `"?zt?rk"`) is data corruption, not conversion: ``` r name <- data.frame(INVNAM = "Öztürk") write_xpt(name, tempfile(fileext = ".xpt"), encoding = "US-ASCII", on_invalid = "translit" ) ``` Error in `write_xpt()`: ! Cannot encode 1 value to "US-ASCII" even after punctuation folding. ✖ Offending value: "Öztürk". ℹ Only smart punctuation has an ASCII fold; use `on_invalid = "fold"` to also strip accents, or write to Dataset-JSON (UTF-8). ## 5. Accent stripping: the full fold, opted into Sometimes stripping accents *is* the documented migration decision — the SAS `BASECHAR()` function exists for exactly this. artoo’s equivalent is `on_invalid = "fold"`: the punctuation table above **plus** the ICU Latin-ASCII transliteration for the rest of the WLATIN1 range (`À–ÿ` to their base letters, `Æ` → `AE`, `ß` → `ss`, `Þ` → `TH`, `×` → `*`, `«»` → `<<` `>>`, `½` → `1/2`). The table is pinned inside artoo, so — unlike `iconv //TRANSLIT`, whose output differs between Linux and macOS — the fold is byte-identical on every platform: ``` r names_df <- data.frame(INVNAM = c("Öztürk", "Straße", "Ærø")) folded_xpt <- tempfile(fileext = ".xpt") write_xpt(names_df, folded_xpt, encoding = "US-ASCII", on_invalid = "fold") ``` Warning in write_xpt(names_df, folded_xpt, encoding = "US-ASCII", on_invalid = "fold"): Folded 3 values to ASCII for "US-ASCII" (accents stripped). ℹ The fold follows the SAS NLS punctuation table plus ICU Latin-ASCII; the original characters are not recoverable from the output. ``` r read_xpt(folded_xpt)$INVNAM ``` [1] "Ozturk" "Strasse" "AEro" The warning is the audit trail: the fold is one-way, and the original characters are not recoverable from the output. A character neither authority maps to ASCII — the Euro sign, the trademark sign, `µ`, `°` — still aborts, because artoo invents no mappings: ``` r cost <- data.frame(COMMENT = "total €100") write_xpt(cost, tempfile(fileext = ".xpt"), encoding = "US-ASCII", on_invalid = "fold" ) ``` Error in `write_xpt()`: ! Cannot encode 1 value to "US-ASCII" even after ASCII folding. ✖ Offending value: "total €100". ℹ The character has no standards-backed ASCII fold (ICU Latin-ASCII leaves it unmapped); write to Dataset-JSON (UTF-8), or set `on_invalid`. The escalation ladder, least to most lossy: `"error"` (refuse) → `"translit"` (fold punctuation only) → `"fold"` (also strip accents) → `"replace"` (`?` per character) → `"ignore"` (drop). Pick the earliest rung your migration plan allows. ## 6. The migration recipe 1. **Read with the true source encoding.** Let [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) detect, or pass `encoding = "wlatin1"` explicitly for formats that cannot carry the answer. Verify with [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) — zero `invalid_encoding` findings means the bytes arrived as characters. 2. **Re-measure lengths in bytes.** Run the `length_overflow` dimension against the UTF-8 form and update the spec where multibyte characters grew. 3. **Write UTF-8 by default** (Dataset-JSON, Parquet, xpt v8 all carry it losslessly). For a US-ASCII submission XPT, write with `encoding = "US-ASCII", on_invalid = "translit"` — punctuation folds, genuine data problems stay loud. Escalate to `on_invalid = "fold"` only when accent stripping is a documented step of the migration plan. ## References - SAS 9.4 NLS Reference Guide, *Migrating Data from WLATIN1 to UTF-8* (the WLATIN1 code-page figures and the smart-punctuation table): - Bouedo, M. (2020), *The SAS Encoding Journey: A Byte at a Time*, SAS Global Forum paper 4561-2020: - FDA Study Data Technical Conformance Guide (the ASCII expectation for submission XPORT): - Unicode Standard Annex \#15, *Unicode Normalization Forms* (the NFC form artoo canonicalises to): # Recipes These are the everyday programming loops a clinical programmer runs — an ADaM analysis dataset, an SDTM domain, a codelist decode, the temporal shapes — each rendered **live** below from the bundled demo data, so what you see is exactly what the code produces. Every recipe ends on a conformed result; the closing line shows how to **ship it** to any deliverable, just by changing the extension. ## An ADaM build: ADSL A real derivation program accumulates working columns the spec never declares — here an age-group cut. Those are “extras”; `extra = "drop"` trims them to exactly the spec’s columns, recorded by the `extra_variable` finding so the drop is never silent. [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) then coerces, orders, sorts, and stamps: ``` r adsl_raw <- cdisc_adsl adsl_raw$AGEGR1_TMP <- cut( adsl_raw$AGE, breaks = c(-Inf, 64, 74, Inf), labels = c("<65", "65-74", ">=75") ) adsl_raw$AGEGR1 <- as.character(adsl_raw$AGEGR1_TMP) adsl <- apply_spec(adsl_raw, adam_spec, "ADSL", extra = "drop") ``` Warning: 1 conformance error for "ADSL". ℹ Run `conformance(x)` on the returned frame to see every finding. ``` r columns(adsl) ``` ADSL -- 43 variables, 60 obs # Variable Type Len Format Label Key 1 STUDYID Char 12 Study Identifier 1 2 USUBJID Char 11 Unique Subject Identifier 2 3 SUBJID Char 4 Subject Identifier for the Study 4 SITEID Char 3 Study Site Identifier 5 SITEGR1 Char 3 Pooled Site Group 1 6 ARM Char 20 Description of Planned Arm 7 TRT01P Char 20 Planned Treatment for Period 01 8 TRT01PN Num Planned Treatment for Period 01 (N) 9 TRT01A Char 20 Actual Treatment for Period 01 10 TRT01AN Num Actual Treatment for Period 01 (N) 11 TRTSDT Num DATE9. Date of First Exposure to Treatment 12 TRTEDT Num DATE9. Date of Last Exposure to Treatment 13 AVGDD Num 5.1 Avg Daily Dose (as planned) 14 CUMDOSE Num 8.1 Cumulative Dose (as planned) 15 AGE Num Age 16 AGEGR1 Char 5 Pooled Age Group 1 17 AGEGR1N Num Pooled Age Group 1 (N) 18 AGEU Char 5 Age Units 19 RACE Char 32 Race 20 RACEN Num Race (N) 21 SEX Char 1 Sex 22 ETHNIC Char 22 Ethnicity 23 SAFFL Char 1 Safety Population Flag 24 ITTFL Char 1 Intent-To-Treat Population Flag 25 EFFFL Char 1 Efficacy Population Flag 26 COMP8FL Char 1 Completers of Week 8 Population Flag 27 COMP16FL Char 1 Completers of Week 16 Population Flag 28 COMP24FL Char 1 Completers of Week 24 Population Flag 29 DISCONFL Char 1 Subject Discontinued Study Flag 30 DSRAEFL Char 1 Subject Discontinued due to AE Flag 31 DTHFL Char 1 Subject Death Flag 32 BMIBL Num 5.1 Baseline BMI (kg/m^2) 33 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 34 HEIGHTBL Num 6.1 Baseline Height (cm) 35 WEIGHTBL Num 6.1 Baseline Weight (kg) 36 EDUCLVL Num Years of Education 37 DURDIS Num 6.1 Duration of Disease (Months) 38 DURDSGR1 Char 4 Pooled Disease Duration Group 1 39 VISIT1DT Num DATE9. Date of Visit 1 40 RFSTDTC Char 10 Subject Reference Start Date/Time 41 RFENDTC Char 10 Subject Reference End Date/Time 42 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) 43 RFENDT Num Date of Discontinuation/Completion ``` r write_xpt(adsl, "adsl.xpt") # ship: or .json / .parquet / .rds ``` ## An SDTM build: DM The pipeline is identical — only the spec and the data change. Assemble the domain (with a QC temporary), conform, then read the result back as a one-line inventory with [`members()`](https://vthanik.github.io/artoo/reference/members.md): ``` r dm_raw <- cdisc_dm dm_raw$AGE_CHECK <- dm_raw$AGE >= 18 dm <- apply_spec(dm_raw, sdtm_spec, "DM", extra = "drop", conformance = "off") json <- tempfile(fileext = ".json") write_json(dm, json) members(json) ``` 1 dataset file member label records variables format file1e9f7d3affea.json DM Demographics 60 15 json ``` r write_xpt(dm, "dm.xpt") # ship: or .json / .parquet / .rds ``` ## Codelists: decode in either direction [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md) reads the codelist a variable is bound to in the spec and maps in either direction — here, deriving the numeric `RACEN` code from the `RACE` decode (`direction = "to_code"`): ``` r adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL", conformance = "off") coded <- decode_column( adsl, adam_spec, "ADSL", from = "RACE", to = "RACEN", direction = "to_code" ) unique(coded[, c("RACE", "RACEN")]) ``` RACE RACEN 1 WHITE 1 20 BLACK OR AFRICAN AMERICAN 2 24 AMERICAN INDIAN OR ALASKA NATIVE 6 A value outside the codelist’s terms aborts (`artoo_error_codelist`) unless you pass `no_match = "keep"` / `"na"`. ## Dates, times, and `--DTC` Clinical dates travel in two shapes, and artoo keeps the distinction explicit through `dataType`, `targetDataType`, and `displayFormat`. **Scope note:** this is about *carriage* — partial-date imputation is an analysis decision your SAP owns; artoo never imputes. An SDTM `--DTC` variable typed `date` is ISO 8601 text by the CDISC storage rule, so it stays character through the pipeline, partial values included (`"2024-03"` is a legal ISO date, and padding it would be imputation by stealth): ``` r dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") class(dm$RFSTDTC) ``` [1] "character" A variable typed `date` with no `targetDataType` realizes to an R `Date` in memory: ``` r adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL", conformance = "off") class(adsl$DISONSDT) ``` [1] "Date" The ADaM numeric-date convention is the other shape: a variable carried as an `integer` SAS-epoch day count, with a `displayFormat` (`date9.`) telling SAS how to render it. `TRTSDT` is one, and the format rides in the metadata: ``` r class(adsl$TRTSDT) ``` [1] "integer" ``` r p <- tempfile(fileext = ".json") write_json(adsl, p) get_meta(read_json(p))@columns$TRTSDT[c("dataType", "displayFormat")] ``` $dataType [1] "integer" $displayFormat [1] "date9." SAS `TIME` values import as `hms` (seconds since midnight), so values past 24 hours, negative values, and fractional seconds all survive: ``` r hms::as_hms(30615) ``` 08:30:15 Either date shape round-trips byte-faithfully — the SAS-epoch (1960-01-01) versus R-epoch conversion happens inside the codecs, never in your code: ``` r identical(read_json(p)$TRTSDT, adsl$TRTSDT) ``` [1] TRUE ## Where to next - [Specifications](https://vthanik.github.io/artoo/articles/specs.md) — the spec each recipe starts from. - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.md) — [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) and the findings in depth. - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.md) — the round trips these ship lines rely on. - [Get started](https://vthanik.github.io/artoo/articles/artoo.md) — the round-trip from the top. # Specifications A `artoo_spec` is artoo’s single source of truth: the variables, CDISC data types, lengths, labels, controlled-terminology codelists, and sort keys for exactly **one** CDISC standard. Read one from the metadata you already have, inspect it as plain data frames, fix it in R when the data disagrees, and write it back — the spec is the contract every later step honors. ## 1. Read a spec [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) ingests a specification from Define-XML 2.x, a Pinnacle 21 workbook, or artoo’s own native JSON, and returns a `artoo_spec`. The bundled ADaM spec also ships as a P21 workbook, so this runs as-is: ``` r p21 <- system.file("extdata", "adam-spec.xlsx", package = "artoo") spec <- read_spec(p21) spec ``` Study: CDISC-Sample Standard: ADaMIG 1.1 Datasets: 2 Variables: 104 Codelists: 30 Methods: 54 Comments: 22 Documents: 9 Spec for: ADSL, ADAE A workbook can carry several standards or duplicate roles; scope the read when you need just one: ``` r read_spec("define.xml", datasets = "ADSL", on_duplicate = "first") ``` ## 2. Inspect with the `spec_*` accessors Each accessor returns a plain data frame (or character vector), so the spec slots straight into ordinary base R work — filter, join, summarise. The datasets a spec covers: ``` r spec_datasets(spec) ``` [1] "ADSL" "ADAE" The variable table is the one you reach for most; here, four columns of it: ``` r spec_variables(spec, "ADSL")[, c("variable", "label", "data_type", "length")] |> head() ``` variable label data_type length 1 STUDYID Study Identifier string 12 2 USUBJID Unique Subject Identifier string 11 3 SUBJID Subject Identifier for the Study string 4 4 SITEID Study Site Identifier string 3 5 SITEGR1 Pooled Site Group 1 string 3 6 ARM Description of Planned Arm string 20 The sort keys that [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) will order by, and the controlled terminology a coded variable is bound to: ``` r spec_keys(spec, "ADSL") ``` [1] "STUDYID" "USUBJID" ``` r head(spec_codelists(spec)) ``` codelist_id order term decode extended comment_id 1 CL.AGEGR1 NA <65 NA 2 CL.AGEGR1 NA 65-80 NA 3 CL.AGEGR1 NA >80 NA 4 CL.AGEGR1N 1 1 <65 NA 5 CL.AGEGR1N 2 2 65-80 NA 6 CL.AGEGR1N 3 3 >80 NA [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md), [`spec_study()`](https://vthanik.github.io/artoo/reference/spec_study.md), [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md), [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md), and [`spec_documents()`](https://vthanik.github.io/artoo/reference/spec_documents.md) expose the remaining slots the same way. ## 3. Fix it in place When the data disagrees with the spec, fix the spec in one line — never reach into internals. [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) retypes a variable; the spec is immutable, so it returns an updated copy: ``` r spec <- set_type(spec, "ADSL", AGE = "float") v <- spec_variables(spec, "ADSL") v$data_type[v$variable == "AGE"] ``` [1] "float" When a check has already found integer-vs-fraction mismatches, [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) applies the fix for every one of them at once, from the findings frame: ``` r findings <- check_spec(cdisc_adsl, spec, "ADSL") spec <- repair_spec(spec, findings) ``` No "integer_fraction" or "integer_overflow" findings to repair. ℹ The spec is returned unchanged. ## 4. Write it back [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) is the inverse of [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) on each format: native JSON is fully lossless, and the P21 workbook is the interchange form. Round-trip a corrected spec through JSON: ``` r out <- tempfile(fileext = ".json") write_spec(spec, out) identical(spec_standard(read_spec(out)), spec_standard(spec)) ``` [1] TRUE Because the two verbs are inverses, format conversion is one composition: ``` r read_spec("define.xml") |> write_spec("spec.xlsx") ``` ## Where to next - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.md) — apply this spec to data, then check every finding. - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.md) — move a conformed dataset between formats without loss. - [Recipes](https://vthanik.github.io/artoo/articles/recipes.md) — the spec in an end-to-end ADaM and SDTM build. - [Get started](https://vthanik.github.io/artoo/articles/artoo.md) — the round-trip from the top. # Authors and Citation ## Authors - **[Vignesh Thanikachalam](https://github.com/vthanik)**. Author, maintainer, copyright holder. ## Citation Source: [`DESCRIPTION`](https://github.com/vthanik/artoo/blob/v0.1.3/DESCRIPTION) Thanikachalam V (2026). *artoo: Lossless CDISC-Native Input and Output for Clinical Datasets*. R package version 0.1.3, . @Manual{, title = {artoo: Lossless CDISC-Native Input and Output for Clinical Datasets}, author = {Vignesh Thanikachalam}, year = {2026}, note = {R package version 0.1.3}, url = {https://vthanik.github.io/artoo/}, } # artoo **artoo** is a lightweight, lossless, CDISC-native reader and writer for clinical-trial datasets. It moves data between **SAS XPORT (XPT)**, **CDISC Dataset-JSON v1.1**, **NDJSON**, **Apache Parquet**, and **RDS** through one canonical metadata model, so converting between any two is lossless *by construction* — not by best effort. ## Installation Install the released version from CRAN: ``` r install.packages("artoo") ``` Or the development version from GitHub: ``` r # install.packages("pak") pak::pak("vthanik/artoo") # or remotes::install_github("vthanik/artoo") ``` ## Quick start A spec describes the dataset; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) conforms a raw frame to it; the writers carry every piece of metadata to disk — one pipeable chain: ``` r library(artoo) # Coerce, order, sort, stamp metadata, then write. The writers return their # input invisibly, so one conformed frame fans out to every deliverable. path <- tempfile(fileext = ".xpt") adsl <- cdisc_adsl |> apply_spec(adam_spec, "ADSL") |> write_xpt(path) #> 6 variables the spec declares are absent from the data (not added): `TRTDURD`, #> `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. #> ℹ See `conformance(x)` for the findings. # Read it back — labels, formats, types, and record count intact. get_meta(read_xpt(path))@dataset$records #> [1] 60 ``` [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) is the quick look a SAS programmer expects from `PROC CONTENTS`, on a conformed frame or straight off a file: ``` r columns(adsl) #> ADSL -- 48 variables, 60 obs #> # Variable Type Len Format Label Key #> 1 STUDYID Char 12 Study Identifier 1 #> 2 USUBJID Char 11 Unique Subject Identifier 2 #> 3 SUBJID Char 4 Subject Identifier for the Study #> 4 SITEID Char 3 Study Site Identifier #> 5 SITEGR1 Char 3 Pooled Site Group 1 #> 6 ARM Char 20 Description of Planned Arm #> 7 TRT01P Char 20 Planned Treatment for Period 01 #> 8 TRT01PN Num Planned Treatment for Period 01 (N) #> 9 TRT01A Char 20 Actual Treatment for Period 01 #> 10 TRT01AN Num Actual Treatment for Period 01 (N) #> 11 TRTSDT Num DATE9. Date of First Exposure to Treatment #> 12 TRTEDT Num DATE9. Date of Last Exposure to Treatment #> 13 AVGDD Num 5.1 Avg Daily Dose (as planned) #> 14 CUMDOSE Num 8.1 Cumulative Dose (as planned) #> 15 AGE Num Age #> 16 AGEGR1 Char 5 Pooled Age Group 1 #> 17 AGEGR1N Num Pooled Age Group 1 (N) #> 18 AGEU Char 5 Age Units #> 19 RACE Char 32 Race #> 20 RACEN Num Race (N) #> 21 SEX Char 1 Sex #> 22 ETHNIC Char 22 Ethnicity #> 23 SAFFL Char 1 Safety Population Flag #> 24 ITTFL Char 1 Intent-To-Treat Population Flag #> 25 EFFFL Char 1 Efficacy Population Flag #> 26 COMP8FL Char 1 Completers of Week 8 Population Flag #> 27 COMP16FL Char 1 Completers of Week 16 Population Flag #> 28 COMP24FL Char 1 Completers of Week 24 Population Flag #> 29 DISCONFL Char 1 Subject Discontinued Study Flag #> 30 DSRAEFL Char 1 Subject Discontinued due to AE Flag #> 31 DTHFL Char 1 Subject Death Flag #> 32 BMIBL Num 5.1 Baseline BMI (kg/m^2) #> 33 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 #> 34 HEIGHTBL Num 6.1 Baseline Height (cm) #> 35 WEIGHTBL Num 6.1 Baseline Weight (kg) #> 36 EDUCLVL Num Years of Education #> 37 DURDIS Num 6.1 Duration of Disease (Months) #> 38 DURDSGR1 Char 4 Pooled Disease Duration Group 1 #> 39 VISIT1DT Num DATE9. Date of Visit 1 #> 40 RFSTDTC Char 10 Subject Reference Start Date/Time #> 41 RFENDTC Char 10 Subject Reference End Date/Time #> 42 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) #> 43 RFENDT Num Date of Discontinuation/Completion #> 44 TRTDUR Num #> 45 DISONSDT Num DATE9. #> 46 DCDECOD Char 27 #> 47 DCREASCD Char 18 #> 48 MMSETOT Num ``` ## Why artoo? - **Lossless by construction.** One canonical metadata model carries labels, CDISC data types, lengths, SAS display formats, controlled-terminology references, and sort keys identically across every format, so any-to-any conversion preserves them — not by best effort, by design. - **Lossless or loud.** A coercion that would truncate or an unencodable byte aborts with a classed condition before it can damage data; there is no silent-truncation path. - **Pure R and lightweight.** No external SAS or Java runtime, and no heavy I/O dependency. - **CDISC-native.** Types, dates and `--DTC` text, and codelists follow the Dataset-JSON v1.1 vocabulary; specs read from Define-XML, Pinnacle 21 workbooks, or native JSON. ## Where artoo fits artoo is the carrier between the formats a clinical-trial dataset travels in: the XPORT a regulator expects, the Dataset-JSON modern CDISC exchange uses, the Parquet an analytics stack reads, and an R-native checkpoint. Reach for it whenever a dataset must change formats without losing the metadata that makes it submission-ready — labels, types, lengths, display formats, codelists, and keys — and you want that guarantee enforced rather than hoped for. It is a focused reader/writer, not a validation suite or a table renderer. ## Supported formats | Format | Reader | Writer | Use | |----|----|----|----| | SAS XPORT (XPT) | [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) | [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) | FDA / PMDA submission | | CDISC Dataset-JSON | [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md) | [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) | Modern CDISC interchange | | NDJSON | [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md) | [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md) | Streaming Dataset-JSON | | Apache Parquet | [`read_parquet()`](https://vthanik.github.io/artoo/reference/read_parquet.md) | [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md) | Analytics, columnar store | | RDS | [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) | [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) | Fast R-native storage | The generic [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) / [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) dispatch on the file extension; every reader supports partial reads via `col_select` and `n_max`. Partial ISO 8601 dates are first-class: a character `--DTC` column typed `date` writes to XPT as ISO text — `"1951-12"` survives byte for byte — while `targetDataType = "integer"` drives the ADaM numeric-date convention. SAS `TIME` values arrive as `hms` (seconds since midnight), and `>24h`, negative, and fractional times round-trip every format. ## Documentation - [Get started](https://vthanik.github.io/artoo/articles/artoo.html) — the whole round-trip, start to finish, on bundled data. - [Specifications](https://vthanik.github.io/artoo/articles/specs.html) — read, inspect, and repair a spec. - [Conform & validate](https://vthanik.github.io/artoo/articles/conform.html) — [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) and every conformance finding. - [Formats & lossless conversion](https://vthanik.github.io/artoo/articles/convert.html) — any-to-any round trips and qualification evidence. - [Recipes](https://vthanik.github.io/artoo/articles/recipes.html) — end-to-end ADaM and SDTM builds, dates, and codelists, rendered live. - [Reference](https://vthanik.github.io/artoo/reference/index.html) — every function, grouped by stage. ## License MIT © Vignesh Thanikachalam # License YEAR: 2026 COPYRIGHT HOLDER: Vignesh Thanikachalam # Changelog ## artoo 0.1.3 CRAN release: 2026-07-22 - [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) gained an `invalid_encoding` dimension (on by default): [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) flags character values whose bytes are not valid UTF-8, the signature of a source read under a mis-declared encoding, before a writer aborts on them. - [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) name resolution now accepts the SAS OEM/DOS encoding names (`pcoem437`, `pcoem850`, `pcoem852`, `pcoem858`, `pcoem862`, `pcoem866`, `msdos737`), and the reference table lists the `PCOEM437` / `PCOEM850` rows. - [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md), [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md), [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md), and [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md) accept `on_invalid = "translit"`, folding smart punctuation (curly quotes, en/em dashes, ellipsis, bullet) to its exact ASCII form per the SAS NLS punctuation table; characters with no fold still abort loudly. - [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md), [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md), [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md), and [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md) also accept `on_invalid = "fold"`: the punctuation fold plus the ICU Latin-ASCII accent strip (`Ö` to `O`, `ß` to `ss`, `Æ` to `AE`), pinned as data so the result is identical on every platform; characters neither table maps (the Euro sign) still abort. - [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) now warns (`artoo_warning_encoding`) when a value forces a column wider than its spec-declared length, instead of widening silently; data is still never truncated. - [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) and the other writers’ `on_invalid = "replace"` now substitutes one `?` per unrepresentable character instead of one per byte (a curly quote previously became `???`). - New article: *Migrating clinical data from WLATIN1 to UTF-8*, including the smart-punctuation fold table and the byte-length migration recipe. ## artoo 0.1.2 CRAN release: 2026-07-02 - Guarded the decimal full-precision JSON round-trip test on `capabilities("long.double")` so it skips on noLD builds, where bit-exact double-to-string round-trips are not guaranteed by the platform C library. ## artoo 0.1.1 CRAN release: 2026-06-24 Initial CRAN release. artoo is a lightweight, lossless, CDISC-native reader and writer for clinical-trial datasets, built around one canonical metadata model (`artoo_meta`) so that conversion between any two supported formats is lossless by construction. Pure R and lightweight, with no external SAS or Java runtime. ### Formats - Reads and writes SAS XPORT (v5 and v8), CDISC Dataset-JSON v1.1, NDJSON, Apache Parquet, and RDS. Every codec carries the full `artoo_meta` — labels, CDISC data types, lengths, SAS display formats, controlled-terminology references, and sort keys — so any-to-any conversion preserves the complete metadata. For Parquet the metadata rides as a `metadata_json` sidecar; a file written by another tool with no sidecar degrades gracefully to a bare frame rather than an error. - Generic [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) / [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) dispatch on the file extension, with [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) / [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) and the matching [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md), [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md), [`read_parquet()`](https://vthanik.github.io/artoo/reference/read_parquet.md), and [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) pairs as direct entry points. Cross-cutting `encoding`, `checks`, and `created` arguments flow through `...`. - Partial reads (`col_select`, `n_max`) on every reader; gzip-transparent JSON and NDJSON; multi-member SAS XPORT libraries via [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md) plus `read_xpt(member = )`. - Numeric fidelity is exact end to end: a `decimal` value is exchanged as a string at IEEE round-trip precision, `integer` values beyond R’s 32-bit range stay numeric rather than overflowing, and `NaN` / infinite values are rejected as invalid CDISC numerics. Rows are sorted in C-locale (byte) order, so a written file is deterministic across locales and matches SAS `PROC SORT` for ASCII keys. - Encodings follow the IANA and SAS standards: the readers and writers accept a charset name in either the SAS or R spelling (see [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md)), character columns are transcoded to UTF-8 and NFC-normalized on read, and the `on_invalid = c("error", "replace", "ignore")` policy governs invalid bytes uniformly across every writer. ### Specifications - [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) builds the canonical metadata model from a Pinnacle 21 Excel workbook, a Define-XML 2.0 / 2.1 file, or a native artoo JSON spec. [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) / [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) dispatch on the file extension: `.xlsx` writes a Pinnacle 21 workbook (Define-XML to P21 is one composition), and the native JSON form is the lossless interchange that round-trips a spec identically. - The spec is single-standard by construction: `@standard` is resolved once from the explicit argument or the source, and study-level fields are canonicalized to the CDISC ODM vocabulary. Accessors include [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md), [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), [`spec_codelists()`](https://vthanik.github.io/artoo/reference/spec_codelists.md), [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md), and [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md). - [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) returns a spec with one or more variables retyped through the CDISC vocabulary; [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) retypes every variable a [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) run flags as fractional or out-of-range under an `integer` data type, so a frame the original spec would refuse coerces after one call. ### Conform and check - `apply_spec(x, spec, dataset, conformance = , na_position = )` coerces each column to its CDISC data type, orders the columns and sorts the rows by the spec’s keys, and stamps the `artoo_meta`. `extra = c("keep", "drop")` controls whether undeclared columns survive; `on_coercion_loss = c("error", "keep")` governs a coercion that would lose data. The pipeline never silently fabricates or drops a column: an undeclared column is reported and kept, a declared-but-absent column is reported and left absent. - [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) validates a data frame against its spec across conformance dimensions toggled by [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md); [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) runs it over a whole study and returns one stacked findings frame; [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md) reads the findings back off a stamped frame. [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) checks a spec for internal consistency against a bundled rule catalog, with no external dependency. - [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md) translates coded values to or from their codelist decodes; [`sync_meta()`](https://vthanik.github.io/artoo/reference/sync_meta.md) reconciles a stamped frame’s metadata after manual edits. ### Inspect - [`members()`](https://vthanik.github.io/artoo/reference/members.md) is the format-neutral inventory of the dataset(s) a path holds, one row per dataset, dispatched through the codec registry. [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) is the SAS PROC CONTENTS / Universal Viewer variable pane over a stamped frame, a plain data frame, or a file path. [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) / [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) read and attach the `artoo_meta`. ### Errors - Every condition artoo raises carries a three-level class chain — `artoo__`, `artoo_`, `artoo_condition` — so a handler can catch a specific kind, a whole severity, or every artoo condition. The data-protection conditions attach their evidence as data (`cnd$variables`, `cnd$findings`) for programmatic inspection. ### Data - Bundled demo specs `adam_spec` (ADaMIG 1.1) and `sdtm_spec` (SDTMIG 3.1.2), built reproducibly from the official CDISC Define-XML 2.1 release examples and shipped also as Pinnacle 21 workbooks under `inst/extdata/`. Demo datasets come from the PHUSE Test Data Factory; the constructor tables `cdisc_adam_datasets` / `cdisc_adam_variables`, `cdisc_sdtm_datasets` / `cdisc_sdtm_variables`, and the shared `cdisc_codelists` build a spec by hand. Every bundled dataset conforms to its bundled spec, gated at build and test time. ### Documentation - An introductory [`vignette("artoo")`](https://vthanik.github.io/artoo/articles/artoo.md) plus task-oriented web articles (specifications; conform and validate; formats and lossless conversion; recipes), and a pkgdown reference site. # Conform a data frame to its spec Run the ordered, transactional artoo pipeline that turns a raw analysis data frame into one conformed to its specification and carrying `artoo_meta`. This is the middle of the workflow (spec -\> apply_spec -\> read\_/write\_): the conformed frame is ready for any `write_*()` codec, and the metadata it now carries makes that write lossless. The input is never mutated; if any step aborts, the call leaves your data untouched. ## Usage ``` r apply_spec( x, spec, dataset, conformance = c("warn", "abort", "off"), na_position = c("first", "last"), extra = c("keep", "drop"), on_coercion_loss = c("error", "keep") ) ``` ## Arguments - x: *The raw data frame to conform.* `: required`. - spec: *The specification to conform to.* `: required`. - dataset: *The dataset whose rules apply.* `: required`. Must name a dataset in `spec`. - conformance: *What to do with conformance findings.* ``. One of: - `"warn"` (default) run [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md), attach the findings (read them with [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md)), warn on any error-severity finding. - `"abort"` abort with `artoo_error_conformance` on any error-severity finding. - `"off"` skip the check entirely. **Note:** this governs only the *findings* disposition — what is *reported*. Pipeline errors are a different category and abort under every setting, including `"off"`: an unknown dataset, and lossy coercion (`artoo_error_type`). Lossy coercion has its own governed gate, `on_coercion_loss`, not this argument; when the spec is the problem, retype it with [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md). - na_position: *Where missing key values sort.* ``. One of `"first"` (default) or `"last"`. `"first"` matches SAS `PROC SORT` (and the FDA submission convention) by ordering missings before present values; `"last"` matches R's [`order()`](https://rdrr.io/r/base/order.html) and the pandas/Polars default. Both are lossless; pick the one your comparison target uses. - extra: *What happens to undeclared columns.* ``. An "extra" is a column of `x` the spec does not declare — typically a derivation temporary. One of: - `"keep"` (default) extras ride along after the declared columns, reported by the `extra_variable` finding. - `"drop"` the returned frame carries exactly the spec's columns; the drop is announced (`artoo_message_apply`), and that message is the audit trail of what was removed. **Interaction:** the drop runs *before* the check, so [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md) reports only the columns the returned frame keeps. Under `conformance = "abort"` an error-severity finding still aborts (those findings arise only on spec-declared columns, which the drop never touches) and the input is never mutated, so the trim cannot mask a failure. **Note:** `"keep"` is the default deliberately. artoo is a lossless carrier, so the metadata step never silently discards a column; extras are surfaced every run (the `extra_variable` finding, and a warning under `conformance = "warn"`), making `"drop"` a conscious opt-in. - on_coercion_loss: *What to do when coercion would lose data.* ``. The governed gate for an `integer` dataType whose data truncates (fractions) or overflows (R's 32-bit range). One of: - `"error"` (default) abort with `artoo_error_type` before any value is touched, refusing to damage the data. - `"keep"` skip coercion for the offending column, leaving it at its wider source type. The values are preserved and the mismatch is reported as an `integer_fraction` / `integer_overflow` finding (read with [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md)), never silently truncated. **Interaction:** independent of `conformance`. `"error"` aborts even under `conformance = "off"`; under `"keep"` the finding it leaves is surfaced by `conformance = "warn"` (the default) and suppressed by `"off"`. **Tip:** `"keep"` is the iterate stance (preserve the data, flag the spec); `"error"` is the submission stance. To fix the spec itself, see [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) and [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md). ## Value *A conformed ``* carrying `artoo_meta` (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)) and, unless `conformance = "off"`, the findings frame [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md) reads back. Hand it to any `write_*()` codec. ## Details **Ordered pipeline.** Four fixed steps run in order: coerce each column to its CDISC dataType, reorder columns to the spec, sort rows by the dataset keys, then stamp the metadata. A spec variable the data lacks is never fabricated as an empty column: artoo is a lossless carrier, not a deriver. It is reported instead, an informational heads-up at apply time plus a `missing_variable` finding (when mandatory) or `missing_permissible` (when not), and left absent, so the conformed frame carries only the columns the data actually had. **Extras are kept by default.** A column the spec does not declare survives the pipeline (ordered after the declared ones), is *reported* by the `extra_variable` conformance finding, and round-trips through every `write_*()` codec with metadata inferred from its R class — membership reported, never enforced by silent destruction. Keeping is the default because artoo is lossless by construction: a metadata-application step that silently discarded columns would break that contract, so trimming data is always an explicit, announced choice rather than a default side effect. `extra = "drop"` opts in to trim-to-spec (the returned frame carries exactly the spec's columns): the undeclared columns are removed *before* the check, so the findings describe exactly the returned frame (a dropped column is never reported as `extra_variable`), and the drop itself is always announced (`artoo_message_apply`) as the audit trail of what was removed — even under `conformance = "off"`. **Lossless or abort, your call.** A coercion that would damage values — an `integer` dataType truncating fractions or overflowing R's 32-bit range — aborts with `artoo_error_type` before any value is touched, under the default `on_coercion_loss = "error"`. This gate is independent of `conformance`: `conformance = "off"` does not bypass it. When the data (not the spec) is right, set `on_coercion_loss = "keep"`: the column keeps its wider source type and the divergence is reported as an `integer_fraction` / `integer_overflow` finding, never silently truncated. When the spec is wrong, retype it with [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) (or [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) from the findings). The error abort carries the offending rows as data: `cnd$variables` is a data frame with columns `variable`, `data_type`, `n`, and `reason` (`"truncated"` / `"overflowed"`), so a pipeline can collect every mismatch in one `tryCatch(..., artoo_error_type = function(cnd) cnd$variables)` pass. The NA-introduction warning (`artoo_warning_coercion`) carries the same frame with `reason = "na_introduced"`, and a `conformance = "abort"` failure carries the complete findings frame as `cnd$findings`. **Values are never translated.** Coded variables keep their submission values (`SEX` stays `"M"`); codelist translation is its own verb, [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md). ## See also **Check:** [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) for the findings; [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md) to read them back. **Fix the spec:** [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) to retype a variable the data disagrees with, [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) to apply every integer fix from a findings frame. **Translate:** [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md) for codelist value mapping. **Metadata:** [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) / [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) for what the stamp attaches. ## Examples ``` r # ---- Example 1: conform ADSL, then read its metadata ---- # # The bundled adam_spec describes ADSL; the raw frame is coerced, # ordered, sorted, and stamped with the CDISC metadata get_meta() reads # back. Variables the spec declares but this extract never derived are # reported (not added), readable via conformance(). adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL") #> 6 variables the spec declares are absent from the data (not added): #> `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. #> ℹ See `conformance(x)` for the findings. get_meta(adsl)@dataset$records #> [1] 60 # ---- Example 2: extras are kept and reported, or dropped on request ---- # # By default a column outside the spec rides along (reported by the # extra_variable finding) and still writes losslessly; extra = "drop" # trims to the spec, announced and still reported. DM is SDTM, so it # conforms against the bundled sdtm_spec. raw <- cdisc_dm raw$DERIVED <- seq_len(nrow(raw)) dm <- apply_spec(raw, sdtm_spec, "DM") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. #> ℹ See `conformance(x)` for the findings. findings <- conformance(dm) findings[findings$check == "extra_variable", c("variable", "message")] #> variable message #> 2 RFXSTDTC Column 'RFXSTDTC' is not declared in the spec. #> 3 RFXENDTC Column 'RFXENDTC' is not declared in the spec. #> 4 RFICDTC Column 'RFICDTC' is not declared in the spec. #> 5 RFPENDTC Column 'RFPENDTC' is not declared in the spec. #> 6 DTHDTC Column 'DTHDTC' is not declared in the spec. #> 7 DTHFL Column 'DTHFL' is not declared in the spec. #> 8 ACTARMCD Column 'ACTARMCD' is not declared in the spec. #> 9 ACTARM Column 'ACTARM' is not declared in the spec. #> 10 DMDTC Column 'DMDTC' is not declared in the spec. #> 11 DMDY Column 'DMDY' is not declared in the spec. #> 12 DERIVED Column 'DERIVED' is not declared in the spec. trimmed <- apply_spec(raw, sdtm_spec, "DM", extra = "drop") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. #> ℹ See `conformance(x)` for the findings. #> Dropped 11 undeclared variables: `RFXSTDTC`, `RFXENDTC`, `RFICDTC`, #> `RFPENDTC`, `DTHDTC`, `DTHFL`, `ACTARMCD`, `ACTARM`, `DMDTC`, `DMDY`, #> and `DERIVED` "DERIVED" %in% names(trimmed) #> [1] FALSE ``` # Control which conformance checks run Build a reusable control that selects which dimensions [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) evaluates. Construct one per study and thread it through every check_spec() call so the conformance surface is consistent. Each toggle is validated at construction, so a mistyped name or value aborts early rather than being silently ignored. ## Usage ``` r artoo_checks( missing_variable = TRUE, missing_permissible = TRUE, extra_variable = TRUE, type_mismatch = TRUE, invalid_encoding = TRUE, length_overflow = TRUE, char_length_limit = TRUE, codelist_membership = TRUE, codelist_membership_extensible = TRUE, label_match = TRUE, key_uniqueness = TRUE, display_format = TRUE, variable_name = TRUE, dataset_name = TRUE, label_length = TRUE, integer_overflow = TRUE, integer_fraction = TRUE, iso8601_format = TRUE ) ``` ## Arguments - missing_variable: *Flag mandatory spec variables absent from the data.* `: default TRUE`. - missing_permissible: *Flag permissible (non-mandatory) spec variables absent from the data.* `: default TRUE`. - extra_variable: *Flag data columns the spec does not declare.* `: default TRUE`. - type_mismatch: *Flag columns whose storage differs from the spec dataType.* `: default TRUE`. - invalid_encoding: *Flag character values whose bytes are not valid UTF-8.* `: default TRUE`. Invalid bytes mean the source was read under a mis-declared encoding; re-read it with the true `encoding=` before the corruption reaches a writer. - length_overflow: *Flag character values longer than the spec length.* `: default TRUE`. - char_length_limit: *Flag character values longer than the SAS XPORT v5 / FDA 200-byte limit.* `: default TRUE`. - codelist_membership: *Flag values outside their closed codelist.* `: default TRUE`. - codelist_membership_extensible: *Flag values outside an extensible codelist's enumerated terms.* `: default TRUE`. A codelist whose `extended` flag is `TRUE` allows sponsor terms, so a non-member is a note, never an error; this toggle silences those notes independently of `codelist_membership`. - label_match: *Flag a column whose label attribute differs from the spec label.* `: default TRUE`. - key_uniqueness: *Flag a dataset whose spec key variables do not uniquely identify its rows.* `: default TRUE`. - display_format: *Flag a date/datetime/time variable whose displayFormat is not a recognized SAS format of that family.* `: default TRUE`. - variable_name: *Flag a data column name that violates the XPORT naming rules.* `: default TRUE`. Over 8 characters (the v5 limit), over 32 (the v8 limit), or containing anything but ASCII letters, digits, and underscore. - dataset_name: *Flag a dataset name that violates the XPORT naming rules.* `: default TRUE`. Same limits as `variable_name`. - label_length: *Flag a column label attribute over the 40-byte XPORT v5 / FDA limit.* `: default TRUE`. - integer_overflow: *Flag an integer-typed variable holding values beyond R's 32-bit integer range.* `: default TRUE`. Such values become `NA` under coercion, so this is an error, not a warning. - integer_fraction: *Flag an integer-typed variable holding fractional values.* `: default TRUE`. Coercion would truncate them (162.6 becomes 162) — a data-integrity event; fix the spec dataType (`float` / `decimal`) or the data before conforming. - iso8601_format: *Flag a character date/datetime/time variable whose values are not valid ISO 8601 text.* `: default TRUE`. A character column under a temporal dataType is the CDISC `--DTC` form; complete values, right-truncated partials (`"1951"`, `"1951-12"`), and SDTMIG hyphen placeholders (`"2003---15"`) all pass, while `"12NOV2019"` or an impossible calendar date is flagged. ## Value *A `` control object*. Pass it as the `checks` argument to [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md). ## Details **Selection, not severity.** This control decides which findings are *produced*; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md)'s `conformance` argument (warn, abort, off) decides what to *do* with the findings its full-default check raises. A disabled dimension is skipped entirely, so the findings frame stays clean. ## See also [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md), which consumes it; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) for the findings disposition. ## Examples ``` r # ---- Example 1: the default runs every conformance dimension ---- # # With no arguments, every conformance dimension is enabled. artoo_checks() #> #> [x] missing_variable #> [x] missing_permissible #> [x] extra_variable #> [x] type_mismatch #> [x] invalid_encoding #> [x] length_overflow #> [x] char_length_limit #> [x] codelist_membership #> [x] codelist_membership_extensible #> [x] label_match #> [x] key_uniqueness #> [x] display_format #> [x] variable_name #> [x] dataset_name #> [x] label_length #> [x] integer_overflow #> [x] integer_fraction #> [x] iso8601_format # ---- Example 2: silence one dimension for a whole study ---- # # Turn off the length check (e.g. while a spec's lengths are provisional) # and reuse the control across every dataset. spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) ck <- artoo_checks(length_overflow = FALSE) nrow(check_spec(cdisc_dm, spec, "DM", checks = ck)) #> [1] 0 ``` # Encodings for clinical datasets, across R, SAS, and Python List the character encodings clinical data actually travels in, with the name each ecosystem uses for the same thing: the R name (the standard IANA name, which [`iconv()`](https://rdrr.io/r/base/iconv.html) and the wider R ecosystem use), the SAS session-encoding name, and the Python codec. Any spelling from the `r` or `sas` column works as the `encoding` argument of every artoo reader and writer. ## Usage ``` r artoo_encodings() ``` ## Value *A ``* with one row per encoding and columns `r` (the R name — the standard IANA name [`iconv()`](https://rdrr.io/r/base/iconv.html) uses, and what artoo records in the metadata), `sas` (the SAS session-encoding name), `python` (the Python codec name), and `description`. ## Details **What an encoding is.** Text is stored as bytes; an encoding is the rule that maps those bytes to characters. Plain A-Z digits and punctuation are the same bytes in every encoding listed here — the differences only show in accented letters (a-umlaut, e-acute), special symbols (micro, degree), and non-Latin scripts. Reading bytes with the wrong rule is what turns a degree sign into garbage. **Which one do I have?** In SAS, run `PROC OPTIONS OPTION=ENCODING; RUN;` and look up the reported name in the `sas` column. Most US/EU Windows SAS installs report `WLATIN1` — that is `windows-1252` here. **Which one should I write?** Usually none: `write_*(encoding = NULL)` (the default) inherits the encoding recorded when the data was read, so a round-trip is byte-faithful. The regulatory defaults artoo applies when nothing is recorded: SAS XPORT writes `US-ASCII` (the FDA Study Data Technical Conformance Guide expectation) and Dataset-JSON / NDJSON write `UTF-8` (required by CDISC and RFC 8259). A value that cannot be represented in the target encoding aborts loudly — see `on_invalid` on the writers. **Note:** in memory, artoo text is always UTF-8 (NFC-normalised) — encodings only matter at the file boundary, exactly as in Python 3. ## See also **Use it:** the `encoding` argument of [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md), [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md), [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md), and the other readers/writers. **Formats:** [`artoo_formats()`](https://vthanik.github.io/artoo/reference/artoo_formats.md) for the codec registry. ## Examples ``` r # ---- Example 1: the full cross-ecosystem table ---- # # One row per encoding; the same byte rule under each ecosystem's name. artoo_encodings() #> r sas python #> 1 UTF-8 UTF-8 utf_8 #> 2 US-ASCII ASCII ascii #> 3 windows-1252 WLATIN1 cp1252 #> 4 windows-1250 WLATIN2 cp1250 #> 5 windows-1251 WCYRILLIC cp1251 #> 6 ISO-8859-1 LATIN1 latin_1 #> 7 ISO-8859-15 LATIN9 iso8859_15 #> 8 CP437 PCOEM437 cp437 #> 9 CP850 PCOEM850 cp850 #> 10 Shift_JIS SHIFT-JIS shift_jis #> 11 CP932 MS-932 cp932 #> 12 EUC-JP EUC-JP euc_jp #> 13 EUC-KR EUC-KR euc_kr #> 14 CP936 MS-936 gbk #> 15 CP950 MS-950 cp950 #> description #> 1 Unicode; Dataset-JSON requirement and the modern default everywhere #> 2 7-bit basic Latin; what the FDA expects inside a submission XPORT #> 3 Western European Windows; the usual US/EU SAS session (WLATIN1) #> 4 Central European Windows #> 5 Cyrillic Windows #> 6 Western European Unix SAS (LATIN1) #> 7 Western European with the euro sign (LATIN9) #> 8 US DOS (OEM); very old PC archives #> 9 Western European DOS (OEM); very old PC archives #> 10 Japanese (PMDA submissions from Japanese SAS sessions) #> 11 Japanese Windows (Microsoft Shift JIS variant) #> 12 Japanese Unix #> 13 Korean #> 14 Simplified Chinese Windows (GBK) #> 15 Traditional Chinese Windows (Big5) # ---- Example 2: look up a SAS session encoding ---- # # PROC OPTIONS reported WLATIN1: find the R and Python names for the # same bytes (the sas and r spellings both work as encoding=). enc <- artoo_encodings() enc[enc$sas == "WLATIN1", ] #> r sas python #> 3 windows-1252 WLATIN1 cp1252 #> description #> 3 Western European Windows; the usual US/EU SAS session (WLATIN1) ``` # Report which formats are available List every registered codec and whether it can read and write in this session. The pure-R formats (xpt, json, rds) are always available; optional-engine formats (parquet) report `FALSE` until their package is installed. Purely informational, modelled on the diagnostic helpers in the wider ecosystem; it never aborts. ## Usage ``` r artoo_formats() ``` ## Value *A ``* with one row per format and columns `format`, `read`, `write` (logical), and `extensions`. ## See also [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) and [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) which use the registry. ## Examples ``` r # ---- Example 1: see what this session can read and write ---- # # rds is always available; the table shows the extensions each codec claims. artoo_formats() #> format read write extensions #> 1 json TRUE TRUE json #> 2 ndjson TRUE TRUE ndjson, jsonl #> 3 parquet TRUE TRUE parquet, pq #> 4 rds TRUE TRUE rds #> 5 xpt TRUE TRUE xpt, xport ``` # Construct a CDISC specification Build and validate a `artoo_spec` from dataset, variable, and codelist tables. Each table is coerced to a plain data frame, missing optional columns are filled with typed `NA`s, every variable type is canonicalised to the CDISC `dataType` vocabulary, and cross-slot integrity (dataset and codelist references) is checked before the object is returned. The spec is the lingua franca the rest of artoo reads, applies, and serialises. ## Usage ``` r artoo_spec( datasets = NULL, variables = NULL, codelists = NULL, study = NULL, values = NULL, methods = NULL, comments = NULL, documents = NULL, standard = NULL ) ``` ## Arguments - datasets: *Dataset-level metadata table.* `: required`. One row per dataset; must carry a `dataset` column. Optional columns `label`, `class`, `structure`, `keys` are filled with `NA` when absent. - variables: *Variable-level metadata table.* `: required`. One row per variable; must carry `dataset`, `variable`, and `data_type`. The `data_type` column is canonicalised to a CDISC `dataType` (e.g. `"text"` becomes `"string"`). **Requirement:** every `dataset` value must appear in `datasets`. - codelists: *Controlled-terminology terms.* ` | NULL`. Must carry `codelist_id` and `term` when supplied. **Interaction:** every `codelist_id` referenced by `variables` must resolve here. - study: *Study-level metadata.* ` | NULL`. A single row of named study fields. Well-known fields are canonicalised to `study_name`, `study_description`, and `protocol_name` (aliases such as `StudyName` or `studyid` resolve automatically); other fields pass through verbatim. A `standard` field, when present, is consumed into `@standard`. - values: *Value-level (VLM) metadata.* ` | NULL`. - methods: *Derivation methods.* ` | NULL`. The Define-XML method definitions variables reference by `method_id`; must carry `method_id` when supplied. Completeness (e.g. a referenced method has a description) is checked by [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md), not here. - comments: *Comment definitions.* ` | NULL`. Referenced by `comment_id`; must carry `comment_id` when supplied. - documents: *Document references.* ` | NULL`. Referenced by `document_id`; must carry `document_id` when supplied. - standard: *The CDISC standard the spec implements.* ` | NULL`. E.g. `"ADaMIG 1.1"` or `"SDTMIG 3.2"`. When `NULL` (default) it is resolved from `datasets$standard` or `study$standard`; absent everywhere, `@standard` is `NA`. **Restriction:** all sources must agree on one value; conflicting standards abort with `artoo_error_spec`. ## Value *A validated `artoo_spec` object.* Inspect it with [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) / [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), or check it with [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## Details **Coerce, then validate.** Each table is first coerced to a plain data frame (a `tibble` is accepted and demoted); known columns are cast to their storage mode and absent optional columns are added as typed `NA`, so every downstream reader can trust the schema. Validation runs only after coercion, on the completed slots. **Type canonicalisation.** `variables$data_type` is mapped through the closed CDISC `dataType` vocabulary (`string`, `integer`, `decimal`, `float`, `double`, `boolean`, `date`, `datetime`, `time`, `URI`). Common SAS / P21 spellings resolve automatically (`"text"`, `"Char"`, `"integer (8)"`, ...); an unrecognised token aborts with `artoo_error_type`. **Cross-slot integrity.** Construction fails (`artoo_error_spec`) if a variable names a dataset absent from `datasets`, or references a `codelist_id` absent from `codelists`. **One spec, one standard.** A `artoo_spec` carries exactly one CDISC standard, stored as the scalar `@standard` property. The constructor resolves it from the `standard` argument, a `standard` column in `datasets` (the P21 workbook shape), and a `standard` field in `study` (the Define-XML shape) — those columns are consumed, so `@standard` is the single home. More than one distinct value aborts with `artoo_error_spec`; scope the source to one standard (e.g. `read_spec(path, datasets = ...)`) instead of mixing. **One study vocabulary.** Well-known study fields are canonicalised to the CDISC ODM GlobalVariables names, snake_cased: `study_name`, `study_description`, `protocol_name`. Source spellings resolve automatically (`StudyName`, `studyid`, ...); fields the vocabulary does not know pass through verbatim. Aliases that disagree on a value abort with `artoo_error_spec`. ## See also **Inspect:** [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md), [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), [`spec_codelists()`](https://vthanik.github.io/artoo/reference/spec_codelists.md), [`spec_keys()`](https://vthanik.github.io/artoo/reference/spec_keys.md), [`spec_study()`](https://vthanik.github.io/artoo/reference/spec_study.md). **Check:** [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). **Predicate:** [`is_artoo_spec()`](https://vthanik.github.io/artoo/reference/is_artoo_spec.md). ## Examples ``` r # ---- Example 1: build a spec from the bundled CDISC-pilot tables ---- # # `cdisc_sdtm_datasets` and `cdisc_sdtm_variables` hold the CDISC pilot SDTM # metadata in the shape artoo_spec() expects; the constructor # canonicalises every type and checks cross-slot integrity. spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) spec_datasets(spec) #> [1] "DM" # ---- Example 2: a focused spec for a single dataset ---- # # Slice the bundled tables to one dataset (DM) to build a smaller spec. dm_ds <- cdisc_sdtm_datasets[cdisc_sdtm_datasets$dataset == "DM", ] dm_var <- cdisc_sdtm_variables[cdisc_sdtm_variables$dataset == "DM", ] dm_spec <- artoo_spec(dm_ds, dm_var, codelists = cdisc_codelists) head(spec_variables(dm_spec, "DM")[, c("variable", "label", "data_type")]) #> variable label data_type #> 1 STUDYID Study Identifier string #> 2 DOMAIN Domain Abbreviation string #> 3 USUBJID Unique Subject Identifier string #> 4 SUBJID Subject Identifier for the Study string #> 5 RFSTDTC Subject Reference Start Date/Time string #> 6 RFENDTC Subject Reference End Date/Time string ``` # artoo: Lossless CDISC-Native Input and Output for Clinical Datasets Reads and writes clinical-trial datasets losslessly across 'SAS' XPORT (XPT), Clinical Data Interchange Standards Consortium (CDISC) Dataset-JSON, and 'Apache Parquet', applying a specification to produce submission-ready Study Data Tabulation Model (SDTM) and Analysis Data Model (ADaM) datasets. A single canonical metadata model carries labels, CDISC data types, lengths, 'SAS' display formats, controlled-terminology references, and sort keys identically across every format, so conversion between any two formats is lossless by construction. Pure 'R' and lightweight, with no external 'SAS' or 'Java' runtime. Implements the published format specifications for CDISC Dataset-JSON () and 'SAS' XPORT (). ## See also Useful links: - - - Report bugs at ## Author **Maintainer**: Vignesh Thanikachalam \[copyright holder\] Authors: - Vignesh Thanikachalam \[copyright holder\] # Demo adverse events analysis dataset (ADaM ADAE) A 60-row sample of the CDISC pilot ADaM adverse events analysis dataset (ADAE): one row per reported event, with treatment-emergent flags, severity, and coding variables (labels preserved as attributes). ## Usage ``` r cdisc_adae ``` ## Format A data frame with 60 rows (`STUDYID`, `USUBJID`, `AETERM`, `AESEV`, `TRTEMFL`, `ASTDT`, ...). ## Source First 60 rows of the CDISC pilot `adam/cdisc/adae.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # Demo subject-level analysis dataset (ADaM ADSL) A 60-subject sample of the CDISC pilot ADaM subject-level analysis dataset (ADSL): one row per subject, with treatment, demographic, baseline, and disposition variables (labels preserved as column attributes). ## Usage ``` r cdisc_adsl ``` ## Format A data frame with 60 rows and 48 variables (`STUDYID`, `USUBJID`, `TRT01P`, `AGE`, `SEX`, `RACE`, `SAFFL`, `TRTSDT`, ...). ## Source First 60 subjects of the CDISC pilot `adam/cdisc/adsl.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # Demo demographics dataset (SDTM DM) A 60-subject sample of the CDISC pilot SDTM demographics domain (DM): one row per subject, with the standard DM variables (labels preserved as attributes). ## Usage ``` r cdisc_dm ``` ## Format A data frame with 60 rows and 25 variables (`STUDYID`, `DOMAIN`, `USUBJID`, `AGE`, `SEX`, `RACE`, `ARM`, `COUNTRY`, ...). ## Source First 60 subjects of the CDISC pilot `sdtm/TDF_SDTM_v1.0/dm.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # CDISC demo specification tables (one standard per pair) The constructor-shaped metadata tables for the bundled demo data, split by CDISC standard because a `artoo_spec` carries exactly one: `cdisc_adam_datasets` + `cdisc_adam_variables` describe ADSL (ADaMIG 1.1), and `cdisc_sdtm_datasets` + `cdisc_sdtm_variables` describe DM (SDTMIG 3.1.2). Each variables table is *derived from the data* (names, labels, inferred CDISC types, byte lengths) by `data-raw/`. Pass one standard's pair to [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md); passing both pairs together aborts with `artoo_error_spec` — mixing standards in one spec is the mistake the split exists to prevent. ## Usage ``` r cdisc_adam_datasets cdisc_adam_variables cdisc_sdtm_datasets cdisc_sdtm_variables cdisc_codelists ``` ## Format Each `*_datasets` table is a data frame with one row per dataset: - dataset: Dataset name (`"ADSL"` or `"DM"`). - label: Dataset label. - standard: The CDISC standard, consumed into the spec's [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md). Each `*_variables` table is a data frame with one row per variable: - dataset: Owning dataset name. - variable: Variable name. - label: Variable label (from the data's `label` attribute). - data_type: CDISC `dataType` inferred from the column's class. - length: Storage length (max byte width for character, 8 for numeric). - order: Variable order within the dataset. - codelist_id: NCI codelist reference (`"C66731"` on `SEX`). `cdisc_codelists` is a data frame of controlled-terminology terms (the real NCI codelist C66731 for `SEX`): - codelist_id: Codelist identifier (`"C66731"`). - term: Submission value (`"M"`, `"F"`, ...). - decode: Decoded value (`"Male"`, `"Female"`, ...). - order: Term order. ## Source Derived from the CDISC pilot `.xpt` files in the public PHUSE Test Data Factory (`phuse-org/phuse-scripts`) by `data-raw/bundle-demo.R`. # Bundled CDISC specifications (ADaM and SDTM) Ready-made `artoo_spec` objects built from the official CDISC Define-XML 2.1 release examples: `adam_spec` (ADaMIG 1.1; datasets ADSL, ADAE) and `sdtm_spec` (SDTMIG 3.1.2; datasets TS, DM, VS, SUPPDM). Every bundled demo dataset conforms to its spec under `apply_spec(conformance = "abort")` — the pairing is gated at build time. The same specs ship as P21 workbooks in `system.file("extdata", "adam-spec.xlsx", package = "artoo")` and `"sdtm-spec.xlsx"`, written by [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md). ## Usage ``` r adam_spec sdtm_spec ``` ## Format A validated [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) object; inspect it with [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md), [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), and [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md). ## Source The CDISC Define-XML 2.1 release example defines (ADaM + SDTM), pinned by sha256 in `data-raw/bundle-spec.R`; data from the PHUSE Test Data Factory. ## Details **Demo adaptations** (each an ADR in `data-raw/bundle-spec.R`): the sponsor-defined codelists (`CL.ARM`, `CL.ARMCD`, `CL.BMICAT`, and the extensible NCI VS codelists) are marked `extended`; `VISITNUM` is typed `float` (the pilot data has fractional visit numbers); VS declares the SDTMIG timepoint variables `VSTPT`/`VSTPTNUM`, with `VSTPTNUM` in the VS key. # Demo supplemental qualifiers dataset (SDTM SUPPDM) A 60-row sample of the CDISC pilot SDTM supplemental qualifiers for DM (SUPPDM): the non-standard qualifier values that ride alongside the DM domain. ## Usage ``` r cdisc_suppdm ``` ## Format A data frame with 60 rows (`STUDYID`, `RDOMAIN`, `USUBJID`, `QNAM`, `QVAL`, ...). ## Source First 60 rows of the CDISC pilot `sdtm/cdiscpilot01/suppdm.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # Demo trial summary dataset (SDTM TS) The CDISC pilot SDTM trial summary domain (TS): one row per trial characteristic (33 rows in the pilot), the study-design parameters a submission carries. ## Usage ``` r cdisc_ts ``` ## Format A data frame with 33 rows (`STUDYID`, `TSPARMCD`, `TSPARM`, `TSVAL`, ...). ## Source The CDISC pilot `sdtm/cdiscpilot01/ts.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # Demo vital signs dataset (SDTM VS) A 60-row sample of the CDISC pilot SDTM vital signs domain (VS): repeated measurements per subject across visits, positions, and planned timepoints. ## Usage ``` r cdisc_vs ``` ## Format A data frame with 60 rows (`STUDYID`, `USUBJID`, `VSTESTCD`, `VSORRES`, `VISITNUM`, `VSPOS`, `VSTPTNUM`, ...). ## Source First 60 rows of the CDISC pilot `sdtm/cdiscpilot01/vs.xpt` from the PHUSE Test Data Factory (`phuse-org/phuse-scripts`). # Check a dataset against its spec Compare a data frame to one dataset's specification and report where they diverge. This is the data-conformance check at the end of the artoo workflow (spec -\> apply_spec -\> check_spec): it reuses the metadata the spec already carries (variables, types, lengths, codelists, keys). It is distinct from [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md), which checks the spec's own internal integrity rather than the data. Both report findings keyed to the same open rule catalog. ## Usage ``` r check_spec( x, spec, dataset, decode = c("none", "to_decode", "to_code"), checks = NULL ) ``` ## Arguments - x: *The data frame to check.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), but any frame works. - spec: *The specification to check against.* `: required`. - dataset: *The dataset whose rules apply.* `: required`. **Restriction:** must name a dataset in `spec` (see [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md)). - decode: *Which codelist column membership is checked against.* ``. One of `"none"` (default), `"to_decode"`, or `"to_code"`. - checks: *Which conformance dimensions to evaluate.* ` | NULL`. When `NULL` (default) every dimension runs; build a control with [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) to disable some. ## Value *A findings data frame* with columns `check`, `dimension`, `severity` (`"error"`, `"warning"`, or `"note"`), `dataset`, `variable`, and `message`, one row per divergence. Zero rows means the data conforms. ## Details **Findings, not enforcement.** `check_spec()` never modifies data; it returns every divergence it finds. [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) runs it and decides what to do via its `conformance` argument (warn, abort, off). The dimensions checked are: missing variables (split into mandatory, an error, and permissible, a warning), extra variables (data column the spec does not declare), type mismatch, ISO 8601 validity of character date/datetime/time values (CDISC partials pass; `"12NOV2019"` does not), fractional values and 32-bit overflow under an `integer` dataType (both would corrupt data at coercion), character length overflow, the hard 200-byte XPORT v5 / FDA character limit, codelist membership, label drift against the spec, key uniqueness, and displayFormat validity. **Decode-aware membership.** `decode` selects which codelist column the data is checked against, matching [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md)'s decode step: `"none"`/`"to_code"` check against the codelist `term`s, `"to_decode"` against the `decode`s. [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) threads its own `decode` through, so a decoded column is not wrongly flagged. **Fatal vs informational coercion checks.** Only `integer_fraction` and `integer_overflow` carry error severity: they mark data an `integer` dataType cannot hold without loss, which [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) refuses to coerce (its `on_coercion_loss` governs that gate). `type_mismatch` is a note: a column stored more widely than the spec declares (an integer-valued double, for instance) coerces cleanly, so it is informational, not a blocker. ## See also [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) which runs this; [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) for the same check across a whole study; [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) to select dimensions; [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) for spec integrity. ## Examples ``` r # ---- Example 1: a conformed frame surfaces only the genuine gaps ---- # # apply_spec() coerces and orders to spec but never fabricates a variable # the data lacks; checking the result reports the permissible variables # this extract never derived (here, six) instead of hiding them as empty # columns. adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL", conformance = "off") #> 6 variables the spec declares are absent from the data (not added): #> `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. nrow(check_spec(adsl, adam_spec, "ADSL")) #> [1] 12 # ---- Example 2: raw data surfaces divergences ---- # # Checking a raw frame with an undeclared column flags the extras. raw <- cdisc_adsl raw$NOTASPEC <- 1 head(check_spec(raw, adam_spec, "ADSL")[, c("check", "variable", "severity")]) #> check variable severity #> 1 missing_permissible TRTDURD warning #> 2 missing_permissible DISONDT warning #> 3 missing_permissible EOSSTT warning #> 4 missing_permissible DCSREAS warning #> 5 missing_permissible EOSDISP warning #> 6 missing_permissible MMS1TSBL warning ``` # Check a whole study against its spec Run [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) over every dataset in a study and return one stacked findings frame. Where [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) answers "does this dataset conform?", `check_study()` answers "is my whole study submittable?" in a single pass, surfacing every dataset's divergences at once instead of one abort at a time. The result is an ordinary findings frame underneath, so filter it by severity or hand it straight to [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md). ## Usage ``` r check_study( spec, data, decode = c("none", "to_decode", "to_code"), checks = NULL ) ``` ## Arguments - spec: *The specification to check against.* `: required`. - data: *The study's datasets.* `: required`. One entry per dataset, named by the dataset (e.g. `list(ADSL = adsl, ADAE = adae)`). Every name must be a dataset in `spec`. - decode: *Which codelist column to check against.* ``. Passed to [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md); one of `"none"` (default), `"to_decode"`, `"to_code"`. - checks: *Which conformance dimensions to run.* ` | NULL`. Passed to [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md); `NULL` (default) runs every dimension. Build a subset with [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md). ## Value *A `` data frame* with the same columns as [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) (`check`, `dimension`, `severity`, `dataset`, `variable`, `message`), one row per divergence across all datasets. Zero rows means the whole study conforms. Print it for the count matrix; treat it as an ordinary data frame otherwise. ## Details **One row per divergence, every dataset stacked.** Each dataset's findings carry its name in the `dataset` column, so the frame is the union of the per-dataset [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) results. Printing renders the dataset-by-check count matrix (the study-level summary); the underlying frame is unchanged. **Data-requiring, like [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md).** `check_study()` checks data against the spec, so it needs the data frames. For the spec's own structural integrity (no data), use [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## See also **One dataset:** [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md). **Spec structure only:** [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). **Repair:** [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) to apply the integer fixes the matrix surfaces. ## Examples ``` r # ---- Example 1: scan a whole study in one pass ---- # # Loop the conformance check over every dataset's data. A fractional AGE # (the spec types it integer) surfaces as an integer_fraction finding; the # print is a dataset-by-check count matrix. adsl <- cdisc_adsl adsl$AGE <- adsl$AGE + 0.5 check_study(adam_spec, list(ADSL = adsl, ADAE = cdisc_adae)) #> 2 datasets: 1 error, 11 warnings, 39 notes #> #> codelist_membership_extensible extra_variable integer_fraction #> ADAE 0 0 0 #> ADSL 1 5 1 #> label_match missing_permissible type_mismatch #> ADAE 7 0 17 #> ADSL 3 6 11 #> #> i Treat this as a findings frame: filter by severity, or pass it to repair_spec(). # ---- Example 2: feed the findings straight into repair_spec() ---- # # The result is an ordinary findings frame, so repair_spec() consumes it to # flip every integer_fraction / integer_overflow variable across the study. findings <- check_study(adam_spec, list(ADSL = adsl)) fixed <- repair_spec(adam_spec, findings) spec_variables(fixed, "ADSL")$data_type[ spec_variables(fixed, "ADSL")$variable == "AGE" ] #> [1] "float" ``` # View a dataset's variable attributes, SAS-style Return a one-row-per-variable attribute table — the pane a SAS programmer reads in `PROC CONTENTS` or the Universal Viewer: position, name, Char/Num type, length, format, informat, label, and the CDISC key sequence. This is the quick look after [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) stamps a frame, or on any dataset file artoo can read. ## Usage ``` r columns(x, member = NULL) ``` ## Arguments - x: *What to describe.* ` | : required`. A stamped frame (carries `artoo_meta`), any plain data frame, or a path to a dataset file (`.xpt`, `.json`, `.ndjson`, `.parquet`, `.rds`). - member: *XPORT member to describe.* ` | NULL`. Only meaningful when `x` is a path to a multi-member `.xpt` file. ## Value *A `` data frame* with columns `#`, `Variable`, `Type`, `Len`, `Format`, `Label`, `Key`, printed left-aligned. The `Informat` column appears only when at least one variable carries an informat (most clinical panes have none). It is an ordinary data frame underneath — filter or inspect it like one. ## Details **Every real column shows.** The table covers the *frame's* columns: a column the spec never declared (which [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) keeps, never drops) still appears, its attributes inferred from the R class. A plain, never-stamped data frame works the same way — every attribute is inferred. **Len is physical storage.** The pane mirrors what `PROC CONTENTS` shows and what the writers store, not a spec digit-width. A Char column always carries a byte `Len` (the declared length, else inferred from the data); a numeric `Len` is blank, because a numeric stores as an 8-byte IEEE double with no character width (a Define-XML numeric `Length` is a digit-width, kept in the metadata for the Define / P21 surface). Format and informat names render uppercase; the metadata keeps the source spelling. **A path reads through the codec.** A file path is dispatched by extension through the same registry as [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md), so the attributes come from the one lossless reader (an unknown extension aborts with the registry's known-extensions message). **Tip:** a multi-member XPORT file needs `member =`; without one the xpt reader aborts and points at [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md) for the listing. **Note:** an `.xpt` path shows a blank `Key`: the XPORT byte layout stores only name, label, length, and formats, so `keySequence` (like codelist and origin) cannot ride in the file. The metadata-carrying formats (`.json`, `.ndjson`, `.parquet`, `.rds`) and the in-session conformed frame show it; re-apply the spec after an xpt read to restore it. ## See also **Members:** [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md) lists a multi-member XPORT file. **Metadata:** [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) for the full `artoo_meta`; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) which stamps it. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: the column pane of a conformed frame ---- # # apply_spec() stamps ADSL with its metadata; columns() reads it back as # the SAS-style attribute table. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") columns(adsl) #> ADSL -- 48 variables, 60 obs #> # Variable Type Len Format Label Key #> 1 STUDYID Char 12 Study Identifier #> 2 USUBJID Char 11 Unique Subject Identifier #> 3 SUBJID Char 4 Subject Identifier for the Study #> 4 SITEID Char 3 Study Site Identifier #> 5 SITEGR1 Char 3 Pooled Site Group 1 #> 6 ARM Char 20 Description of Planned Arm #> 7 TRT01P Char 20 Planned Treatment for Period 01 #> 8 TRT01PN Num Planned Treatment for Period 01 (N) #> 9 TRT01A Char 20 Actual Treatment for Period 01 #> 10 TRT01AN Num Actual Treatment for Period 01 (N) #> 11 TRTSDT Num Date of First Exposure to Treatment #> 12 TRTEDT Num Date of Last Exposure to Treatment #> 13 TRTDUR Num Duration of Treatment (days) #> 14 AVGDD Num Avg Daily Dose (as planned) #> 15 CUMDOSE Num Cumulative Dose (as planned) #> 16 AGE Num Age #> 17 AGEGR1 Char 5 Pooled Age Group 1 #> 18 AGEGR1N Num Pooled Age Group 1 (N) #> 19 AGEU Char 5 Age Units #> 20 RACE Char 32 Race #> 21 RACEN Num Race (N) #> 22 SEX Char 1 Sex #> 23 ETHNIC Char 22 Ethnicity #> 24 SAFFL Char 1 Safety Population Flag #> 25 ITTFL Char 1 Intent-To-Treat Population Flag #> 26 EFFFL Char 1 Efficacy Population Flag #> 27 COMP8FL Char 1 Completers of Week 8 Population Flag #> 28 COMP16FL Char 1 Completers of Week 16 Population Flag #> 29 COMP24FL Char 1 Completers of Week 24 Population Flag #> 30 DISCONFL Char 1 Did the Subject Discontinue the Study? #> 31 DSRAEFL Char 1 Discontinued due to AE? #> 32 DTHFL Char 1 Subject Died? #> 33 BMIBL Num Baseline BMI (kg/m^2) #> 34 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 #> 35 HEIGHTBL Num Baseline Height (cm) #> 36 WEIGHTBL Num Baseline Weight (kg) #> 37 EDUCLVL Num Years of Education #> 38 DISONSDT Num Date of Onset of Disease #> 39 DURDIS Num Duration of Disease (Months) #> 40 DURDSGR1 Char 4 Pooled Disease Duration Group 1 #> 41 VISIT1DT Num Date of Visit 1 #> 42 RFSTDTC Char 10 Subject Reference Start Date/Time #> 43 RFENDTC Char 10 Subject Reference End Date/Time #> 44 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) #> 45 RFENDT Num Date of Discontinuation/Completion #> 46 DCDECOD Char 27 Standardized Disposition Term #> 47 DCREASCD Char 18 Reason for Discontinuation #> 48 MMSETOT Num MMSE Total # ---- Example 2: straight off a file ---- # # Write the conformed frame to any format and point columns() at the # path; the codec reads it back and the attributes are identical. p <- tempfile(fileext = ".json") write_json(adsl, p) columns(p) #> ADSL -- 48 variables, 60 obs #> # Variable Type Len Format Label Key #> 1 STUDYID Char 12 Study Identifier #> 2 USUBJID Char 11 Unique Subject Identifier #> 3 SUBJID Char 4 Subject Identifier for the Study #> 4 SITEID Char 3 Study Site Identifier #> 5 SITEGR1 Char 3 Pooled Site Group 1 #> 6 ARM Char 20 Description of Planned Arm #> 7 TRT01P Char 20 Planned Treatment for Period 01 #> 8 TRT01PN Num Planned Treatment for Period 01 (N) #> 9 TRT01A Char 20 Actual Treatment for Period 01 #> 10 TRT01AN Num Actual Treatment for Period 01 (N) #> 11 TRTSDT Num Date of First Exposure to Treatment #> 12 TRTEDT Num Date of Last Exposure to Treatment #> 13 TRTDUR Num Duration of Treatment (days) #> 14 AVGDD Num Avg Daily Dose (as planned) #> 15 CUMDOSE Num Cumulative Dose (as planned) #> 16 AGE Num Age #> 17 AGEGR1 Char 5 Pooled Age Group 1 #> 18 AGEGR1N Num Pooled Age Group 1 (N) #> 19 AGEU Char 5 Age Units #> 20 RACE Char 32 Race #> 21 RACEN Num Race (N) #> 22 SEX Char 1 Sex #> 23 ETHNIC Char 22 Ethnicity #> 24 SAFFL Char 1 Safety Population Flag #> 25 ITTFL Char 1 Intent-To-Treat Population Flag #> 26 EFFFL Char 1 Efficacy Population Flag #> 27 COMP8FL Char 1 Completers of Week 8 Population Flag #> 28 COMP16FL Char 1 Completers of Week 16 Population Flag #> 29 COMP24FL Char 1 Completers of Week 24 Population Flag #> 30 DISCONFL Char 1 Did the Subject Discontinue the Study? #> 31 DSRAEFL Char 1 Discontinued due to AE? #> 32 DTHFL Char 1 Subject Died? #> 33 BMIBL Num Baseline BMI (kg/m^2) #> 34 BMIBLGR1 Char 6 Pooled Baseline BMI Group 1 #> 35 HEIGHTBL Num Baseline Height (cm) #> 36 WEIGHTBL Num Baseline Weight (kg) #> 37 EDUCLVL Num Years of Education #> 38 DISONSDT Num Date of Onset of Disease #> 39 DURDIS Num Duration of Disease (Months) #> 40 DURDSGR1 Char 4 Pooled Disease Duration Group 1 #> 41 VISIT1DT Num Date of Visit 1 #> 42 RFSTDTC Char 10 Subject Reference Start Date/Time #> 43 RFENDTC Char 10 Subject Reference End Date/Time #> 44 VISNUMEN Num End of Trt Visit (Vis 12 or Early Term.) #> 45 RFENDT Num Date of Discontinuation/Completion #> 46 DCDECOD Char 27 Standardized Disposition Term #> 47 DCREASCD Char 18 Reason for Discontinuation #> 48 MMSETOT Num MMSE Total ``` # Read the conformance findings a dataset carries Pull the conformance findings [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) attached to a conformed data frame — the readable answer to "what did the check find?". The result is the same findings frame [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) returns (one row per divergence), with a print method that renders a sectioned report, so `conformance(adsl)` at the console is the inspection step the `artoo_warning_conformance` warning points you at. ## Usage ``` r conformance(x) ``` ## Arguments - x: *A data frame produced by [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md).* `: required`. **Requirement:** the conformance check must have run: a frame from `apply_spec(..., conformance = "off")` (or one rebuilt by a transform that dropped attributes) carries no findings and aborts with `artoo_error_input`. ## Value *A `` data frame* with columns `check`, `dimension`, `severity` (`"error"`, `"warning"`, or `"note"`), `dataset`, `variable`, and `message`. Zero rows means the data conformed. Print it for the sectioned report; treat it as an ordinary data frame for programmatic use. ## See also [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) which attaches the findings; [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) for the same check on demand; [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) to select dimensions. ## Examples ``` r spec <- artoo_spec( cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists ) # ---- Example 1: inspect what the conform step found ---- # # Conforming raw DM records the findings on the result; conformance() # renders them as a report instead of a raw attribute. dm <- suppressWarnings(apply_spec(cdisc_dm, spec, "DM")) conformance(dm) #> : 0 errors, 0 warnings, 0 notes #> No findings. The data conforms to the spec. # ---- Example 2: gate a pipeline on error-severity findings ---- # # The findings frame is an ordinary data frame: filter by severity to # drive your own logic. f <- conformance(dm) nrow(f[f$severity == "error", ]) #> [1] 0 ``` # Derive or translate a variable through its codelist Map one column's values through a spec codelist — code to decode or decode to code — writing the result to a new variable or in place. This is the everyday companion to [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md)'s whole-dataset `decode` step: deriving `RACEN` from `RACE`, recovering submission codes from decoded values, or decoding a single variable for display, without re-running the pipeline. When the target variable is declared in the spec, the result is also coerced to its dataType and labelled, so the new column lands conformed. ## Usage ``` r decode_column( x, spec, dataset, from, to = from, direction = c("to_decode", "to_code"), no_match = c("error", "keep", "na"), trim = TRUE, ignore_case = FALSE ) ``` ## Arguments - x: *The data frame to extend.* `: required`. - spec: *The specification carrying the codelists.* `: required`. - dataset: *The dataset whose variables apply.* `: required`. Must name a dataset in `spec`. - from: *The source column.* `: required`. Must be a column of `x`. - to: *The destination variable.* `: default `from“. Defaults to translating in place. A `to` declared in the spec gets its dataType coercion and label; an undeclared \`to\` is plain character. - direction: *Which way to map.* ``. One of: - `"to_decode"` (default) map codes to their decoded values (`"M"` becomes `"Male"`). - `"to_code"` map decoded values to their submission codes — the `RACEN`-from-`RACE` derivation. - no_match: *Policy for values absent from the codelist.* ``. One of `"error"` (default; abort with `artoo_error_codelist`, naming the unmatched values and the codelist), `"keep"` (carry the source value through), or `"na"`. - trim: *Match after trimming whitespace.* `: default TRUE`. - ignore_case: *Match case-insensitively.* `: default FALSE`. Case differences are usually genuine CT violations, so this is opt-in. ## Value *The data frame `x`* with the `to` column added (at the end) or replaced (in place), ready for the next pipeline step. ## Details **Which codelist applies.** The codelist attached to `to` in the spec wins (the natural direction for `RACEN`-style derivations, where the numeric variable owns the code/decode pairs); when `to` declares none, `from`'s codelist is used. If neither variable references a codelist the call aborts — there is nothing to map through. **Mismatched surfaces chain.** A single call maps through ONE codelist, so the winning codelist's terms (or decodes) must line up with the `from` values — the CDISC `*N` convention guarantees this for `RACEN`-style pairs, whose decodes are the character variable's submission values. When the two codelists share no value surface (say `SEXN`'s decodes are `"Female"`/`"Male"` but `SEX` holds `"F"`/`"M"`), the unmatched values hit the `no_match` policy; translate in two hops instead — decode through `from`'s codelist first, then `to_code` through the destination's: dm |> decode_column(spec, "DM", from = "SEX", to = "SEXDECD") |> decode_column(spec, "DM", from = "SEXDECD", to = "SEXN", direction = "to_code") **Soft matches are reported, never silent.** Values that match only after trimming whitespace (or case-folding, when `ignore_case = TRUE`) still map, with a `artoo_warning_codelist` naming the variants — [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) always compares exactly, so clean the source for submission. ## See also **Whole-dataset decode:** [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) with `decode =`. **Inspect the terms:** [`spec_codelists()`](https://vthanik.github.io/artoo/reference/spec_codelists.md). **Check membership:** [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md). ## Examples ``` r spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) # ---- Example 1: decode a coded variable into a display column ---- # # SEX is coded against C66731; map the codes to their decodes in a new # column, leaving the submission values untouched. dm <- decode_column(cdisc_dm, spec, "DM", from = "SEX", to = "SEXDECD") table(dm$SEX, dm$SEXDECD) #> #> Female Male #> F 29 0 #> M 0 31 # ---- Example 2: the RACEN pattern, a coded numeric from its decode ---- # # Declare SEXN as an integer variable owning a numeric codelist, then # derive it from SEX's decoded values: to_code maps each decode to its # submission code, and the spec dataType makes the result integer. vars <- rbind( cdisc_sdtm_variables, data.frame( dataset = "DM", variable = "SEXN", label = "Sex (N)", data_type = "integer", length = 8L, order = NA_integer_, codelist_id = "SEXN" ) ) cls <- rbind( cdisc_codelists, data.frame( codelist_id = "SEXN", term = c("1", "2"), decode = c("F", "M"), order = 1:2 ) ) spec_n <- artoo_spec(cdisc_sdtm_datasets, vars, codelists = cls) dm_n <- decode_column(cdisc_dm, spec_n, "DM", from = "SEX", to = "SEXN", direction = "to_code" ) str(dm_n$SEXN) #> int [1:60] 1 2 2 2 1 1 1 2 1 2 ... #> - attr(*, "label")= chr "Sex (N)" ``` # Read the metadata a dataset carries Pull the `artoo_meta` off a data frame produced by [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) or read back by any `read_*()` codec. The metadata travels as a single Dataset-JSON string in the frame's `metadata_json` attribute; `get_meta()` parses it to the S7 object, the form every codec writes from. This is the read half of the lossless round-trip. ## Usage ``` r get_meta(x) ``` ## Arguments - x: *A data frame carrying artoo metadata.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) or a `read_*()` codec. **Requirement:** `x` must carry a `metadata_json` attribute (set by [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md), [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), or a reader); a bare frame aborts with `artoo_error_input`. ## Value *A ``* with two properties. `@dataset` is a named list of dataset-level attributes: `itemGroupOID`, `name`, `label`, `records`, `studyOID`, `metaDataVersionOID`, `encoding`, and `keys`. `@columns` is a named list with one entry per variable, each carrying `itemOID`, `name`, `label`, `dataType`, `targetDataType`, `length`, `displayFormat`, `informat`, `keySequence`, `codelist`, `significantDigits`, and `origin` (absent values are `NULL`). Pass it to [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) to re-attach, or index it directly (`meta@columns$AGE$label`). ## See also [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) for the write half; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) which stamps it. ## Examples ``` r # ---- Example 1: read metadata off a conformed dataset ---- # # apply_spec() stamps the metadata; get_meta() reads it back as the S7 # object whose @columns holds one CDISC attribute set per variable. spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) adsl <- apply_spec(cdisc_adsl, spec, "ADSL") meta <- get_meta(adsl) meta@columns$STUDYID #> $itemOID #> [1] "IT.ADSL.STUDYID" #> #> $name #> [1] "STUDYID" #> #> $label #> [1] "Study Identifier" #> #> $dataType #> [1] "string" #> #> $length #> [1] 12 #> # ---- Example 2: round-trip metadata across two frames ---- # # The metadata is a portable object: read it off one frame and stamp it # onto another with set_meta(). bare <- as.data.frame(adsl) attr(bare, "metadata_json") <- NULL restamped <- set_meta(bare, meta) identical(get_meta(restamped)@columns, meta@columns) #> [1] TRUE ``` # Package index ## Specs Build a artoo_spec — the canonical CDISC-shaped description of your datasets, one CDISC standard each — or read one from native JSON, a Pinnacle 21 workbook, or Define-XML, and write it back out. Amend it in R when the data disagrees, then read any slot back with the spec\_\* accessors. - [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) : Construct a CDISC specification - [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) : Read a specification from JSON, Excel, or Define-XML - [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) : Write a specification to native JSON or a P21 Excel workbook - [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) : Override a variable's dataType in a spec - [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) : Repair a spec from its conformance findings - [`is_artoo_spec()`](https://vthanik.github.io/artoo/reference/is_artoo_spec.md) : Test for a artoo_spec object - [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md) : The CDISC standard a spec implements - [`spec_study()`](https://vthanik.github.io/artoo/reference/spec_study.md) : Study-level metadata - [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) : Dataset names in a spec - [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md) : Variables in a spec - [`spec_codelists()`](https://vthanik.github.io/artoo/reference/spec_codelists.md) : Codelist terms - [`spec_keys()`](https://vthanik.github.io/artoo/reference/spec_keys.md) : Sort keys for a dataset - [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md) : Derivation methods in a spec - [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md) : Comment definitions in a spec - [`spec_documents()`](https://vthanik.github.io/artoo/reference/spec_documents.md) : Document references in a spec ## Conform & validate Apply the spec to a raw frame — coerce, order, sort, stamp metadata — decode single variables through its codelists, and read or replace the artoo_meta the result carries. Then surface every conformance finding for one dataset or a whole study, plus the spec’s own integrity, with the control object that scopes both. - [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) : Conform a data frame to its spec - [`decode_column()`](https://vthanik.github.io/artoo/reference/decode_column.md) : Derive or translate a variable through its codelist - [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) : Read the metadata a dataset carries - [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) : Attach metadata to a dataset - [`sync_meta()`](https://vthanik.github.io/artoo/reference/sync_meta.md) : Re-align metadata with a transformed data frame - [`is_artoo_meta()`](https://vthanik.github.io/artoo/reference/is_artoo_meta.md) : Test for a artoo_meta object - [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) : Check a dataset against its spec - [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) : Check a whole study against its spec - [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) : Validate a specification for submission-readiness - [`conformance()`](https://vthanik.github.io/artoo/reference/conformance.md) : Read the conformance findings a dataset carries - [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) : Control which conformance checks run - [`is_artoo_checks()`](https://vthanik.github.io/artoo/reference/is_artoo_checks.md) : Test for a artoo_checks control ## Read and write Lossless dataset I/O across every supported format — generic dispatch on the file extension, plus a short wrapper per format — and the SAS-viewer-style variable pane and dataset inventory for any file. - [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) : Read a dataset from any supported format - [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) : Write a dataset to any supported format - [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) : Read a dataset from SAS XPORT - [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) : Write a dataset to SAS XPORT - [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md) : Read a dataset from CDISC Dataset-JSON - [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) : Write a dataset to CDISC Dataset-JSON - [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md) : Read a dataset from CDISC Dataset-JSON NDJSON - [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md) : Write a dataset to CDISC Dataset-JSON NDJSON - [`read_parquet()`](https://vthanik.github.io/artoo/reference/read_parquet.md) : Read a dataset from Apache Parquet - [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md) : Write a dataset to Apache Parquet - [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) : Read a dataset from rds - [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) : Write a dataset to rds - [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) : View a dataset's variable attributes, SAS-style - [`members()`](https://vthanik.github.io/artoo/reference/members.md) : List the datasets in a file or directory - [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md) : List the members of a SAS XPORT transport file ## Reference data Reference tables for the codecs this session can read and write and the encoding names R, SAS, and Python share, plus the bundled CDISC pilot specs, metadata tables, and datasets used throughout the docs — all rebuilt from public sources. - [`artoo_formats()`](https://vthanik.github.io/artoo/reference/artoo_formats.md) : Report which formats are available - [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) : Encodings for clinical datasets, across R, SAS, and Python - [`adam_spec`](https://vthanik.github.io/artoo/reference/cdisc_specs.md) [`sdtm_spec`](https://vthanik.github.io/artoo/reference/cdisc_specs.md) : Bundled CDISC specifications (ADaM and SDTM) - [`cdisc_adam_datasets`](https://vthanik.github.io/artoo/reference/cdisc_spec.md) [`cdisc_adam_variables`](https://vthanik.github.io/artoo/reference/cdisc_spec.md) [`cdisc_sdtm_datasets`](https://vthanik.github.io/artoo/reference/cdisc_spec.md) [`cdisc_sdtm_variables`](https://vthanik.github.io/artoo/reference/cdisc_spec.md) [`cdisc_codelists`](https://vthanik.github.io/artoo/reference/cdisc_spec.md) : CDISC demo specification tables (one standard per pair) - [`cdisc_adsl`](https://vthanik.github.io/artoo/reference/cdisc_adsl.md) : Demo subject-level analysis dataset (ADaM ADSL) - [`cdisc_adae`](https://vthanik.github.io/artoo/reference/cdisc_adae.md) : Demo adverse events analysis dataset (ADaM ADAE) - [`cdisc_dm`](https://vthanik.github.io/artoo/reference/cdisc_dm.md) : Demo demographics dataset (SDTM DM) - [`cdisc_vs`](https://vthanik.github.io/artoo/reference/cdisc_vs.md) : Demo vital signs dataset (SDTM VS) - [`cdisc_ts`](https://vthanik.github.io/artoo/reference/cdisc_ts.md) : Demo trial summary dataset (SDTM TS) - [`cdisc_suppdm`](https://vthanik.github.io/artoo/reference/cdisc_suppdm.md) : Demo supplemental qualifiers dataset (SDTM SUPPDM) # Test for a artoo_checks control Report whether an object is a `artoo_checks` control built by [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md). Use it to guard a `checks` argument before threading it into [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) or [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md). ## Usage ``` r is_artoo_checks(x) ``` ## Arguments - x: *Object to test.* ``. ## Value *A ``*: `TRUE` when `x` is a `artoo_checks`. ## See also [`artoo_checks()`](https://vthanik.github.io/artoo/reference/artoo_checks.md) to build one. ## Examples ``` r # ---- Example 1: confirm a control before reusing it ---- # # is_artoo_checks() distinguishes a real control from a bare list of flags. is_artoo_checks(artoo_checks()) #> [1] TRUE is_artoo_checks(list(missing_variable = TRUE)) #> [1] FALSE ``` # Test for a artoo_meta object Report whether an object is a `artoo_meta` — the CDISC-shaped metadata a conformed dataset carries through the artoo workflow (spec -\> apply_spec -\> read\_/write\_). [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) returns one; this is the type guard before you inspect its `@dataset` and `@columns` slots. ## Usage ``` r is_artoo_meta(x) ``` ## Arguments - x: *Object to test.* ``. ## Value *A ``*: `TRUE` when `x` is a `artoo_meta`, else `FALSE`. ## See also [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) and [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md) to read and attach metadata. ## Examples ``` r # ---- Example 1: guard before inspecting metadata ---- # # get_meta() yields a artoo_meta; is_artoo_meta() confirms the type before # you reach into its slots. spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) adsl <- apply_spec(cdisc_adsl, spec, "ADSL") meta <- get_meta(adsl) is_artoo_meta(meta) #> [1] TRUE # ---- Example 2: a bare data frame carries no meta object ---- # # The raw frame itself is not a artoo_meta — only the object get_meta() # returns is. is_artoo_meta(cdisc_adsl) #> [1] FALSE ``` # Test for a artoo_spec object Report whether an object is a `artoo_spec` — the validated CDISC specification that drives the artoo workflow (spec -\> apply_spec -\> read\_/write\_). [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) builds one; this is the type guard before you pass it to [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) or reach into it with the spec accessors. ## Usage ``` r is_artoo_spec(x) ``` ## Arguments - x: *Object to test.* ``. ## Value *A ``*: `TRUE` when `x` is a `artoo_spec`, else `FALSE`. ## See also [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) to build one; [`is_artoo_meta()`](https://vthanik.github.io/artoo/reference/is_artoo_meta.md) for the metadata guard. ## Examples ``` r # ---- Example 1: guard a built specification ---- # # artoo_spec() assembles and validates a spec; is_artoo_spec() confirms the # type before you drive apply_spec() with it. spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) is_artoo_spec(spec) #> [1] TRUE # ---- Example 2: an ordinary object is not a spec ---- # # Any non-artoo_spec value — a bare data frame, say — returns FALSE. is_artoo_spec(cdisc_dm) #> [1] FALSE ``` # List the datasets in a file or directory Inventory the dataset(s) a path contains, one row per dataset, dispatched by extension through the same codec registry as [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md). A SAS XPORT library lists every member; a single-dataset file (`.json`, `.ndjson`, `.parquet`, `.rds`) reports one row; a directory inventories each dataset file it holds. The format-neutral companion to the xpt-specific [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md). ## Usage ``` r members(path) ``` ## Arguments - path: *A dataset file or a directory.* `: required`. A path to a dataset file (`.xpt`, `.json`, `.ndjson`, `.parquet`, `.rds`) or to a directory holding such files. A path that does not exist, or a file whose extension no codec claims, aborts. ## Value *A `` data frame*, one row per dataset, with columns `file` (source basename), `member` (dataset name), `label`, `records` (row count), `variables` (column count), and `format` (the codec format). Empty when a directory holds no dataset files. It is an ordinary data frame underneath. ## Details **One dataset per file, except XPORT.** XPORT is the only multi-dataset container artoo handles, so only an `.xpt` path can return more than one row. Every other format is one dataset per file. **A directory is inventoried, not descended.** Only the files directly in the directory are listed (no recursion); files whose extension no codec claims are skipped, and a directory with no dataset files returns an empty inventory rather than aborting. A dataset file that fails to read aborts with its codec's error, naming the file. **Note:** counting `records` reads the file through its codec (the one lossless reader), so members() is an honest count, not a header guess; for a large directory it reads every dataset. ## See also **Members of one XPORT file:** [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md). **Per-variable attributes:** [`columns()`](https://vthanik.github.io/artoo/reference/columns.md) for one dataset's variable pane. ## Examples ``` r dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. # ---- Example 1: one dataset in a file ---- # # A single-dataset format reports exactly one member. p <- tempfile(fileext = ".json") write_json(dm, p) members(p) #> 1 dataset #> file member label records variables format #> file1a019a22425.json DM Demographics 60 25 json # ---- Example 2: every dataset in a directory ---- # # Point members() at a folder to inventory each dataset file it holds, one # row per dataset, dispatched by extension. dir <- tempfile("datasets") dir.create(dir) write_json(dm, file.path(dir, "dm.json")) write_rds(dm, file.path(dir, "dm.rds")) members(dir) #> 2 datasets #> file member label records variables format #> dm.json DM Demographics 60 25 json #> dm.rds DM Demographics 60 25 rds ``` # Read a dataset from any supported format Read a clinical file back to a data frame, restoring its `artoo_meta`. The codec is chosen from the file extension (or an explicit `format`), and the metadata the file carries is re-attached, so a value written by [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) round-trips losslessly. This is the ingest end of the I/O layer; the per-format wrappers like [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) call it. ## Usage ``` r read_dataset(path, format = NULL, col_select = NULL, n_max = Inf, ...) ``` ## Arguments - path: *Source file path.* `: required`. Its extension selects the codec unless `format` is given. - format: *Force a codec instead of inferring from the extension.* ` | NULL`. One of the registered formats (see [`artoo_formats()`](https://vthanik.github.io/artoo/reference/artoo_formats.md)). - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names. Columns return in file order (not the requested order) and the `artoo_meta` is filtered to match. Works on every format: parquet narrows columns natively, the rest filter after decode. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. xpt v8 bounds the disk read; the other formats cap after decode. - ...: *Codec-specific arguments* passed through to the decoder (see the per-format wrappers, e.g. [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md)). An argument the codec does not know is an error, never silently ignored. ## Value *A ``* carrying `artoo_meta` when the file recorded it (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)). A file whose payload is not a data frame is a `artoo_error_codec`. ## See also [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the inverse; [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) for the per-format wrapper. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: round-trip a dataset through rds ---- # # Write a conformed dataset, then read it back; the metadata survives. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".rds") write_dataset(adsl, path) back <- read_dataset(path) identical(get_meta(back)@columns, get_meta(adsl)@columns) #> [1] TRUE # ---- Example 2: the metadata names the dataset and row count ---- # # The restored artoo_meta exposes the dataset-level attributes. get_meta(back)@dataset$records #> [1] 60 ``` # Read a dataset from CDISC Dataset-JSON Read a CDISC Dataset-JSON v1.1 (`.json`) file back to a data frame, restoring the full `artoo_meta` it carries and realizing SAS date/datetime/time variables to R `Date` / `POSIXct` / [`hms::hms`](https://hms.tidyverse.org/reference/hms.html). Column types are reconstructed from the recorded metadata, not guessed from the JSON tokens, so the round-trip is lossless. The ingest end of the I/O layer; a thin wrapper over [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) with `format = "json"`. ## Usage ``` r read_json(path, col_select = NULL, n_max = Inf, encoding = NULL) ``` ## Arguments - path: *Source `.json` path.* `: required`. A JSON file that is not Dataset-JSON v1.1 aborts with `artoo_error_codec`. - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names. Columns return in file order (not the requested order) and the `artoo_meta` is filtered to match. Works on every format: parquet narrows columns natively, the rest filter after decode. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. xpt v8 bounds the disk read; the other formats cap after decode. - encoding: *Source charset of the file bytes.* ` | NULL`. `NULL` (default) reads UTF-8, as Dataset-JSON requires. Pass an IANA or SAS charset name (e.g. `"windows-1252"`) only to read a non-conformant file a producer wrote in that charset; the bytes are transcoded to UTF-8 on read. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. ## Value *A ``* carrying `artoo_meta` (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)). ## See also [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) for the inverse; [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: round-trip a conformed dataset through Dataset-JSON ---- # # The variable labels, types, and keys survive the round-trip. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".json") write_json(adsl, path) back <- read_json(path) identical(get_meta(back)@columns, get_meta(adsl)@columns) #> [1] TRUE # ---- Example 2: the metadata names the dataset and row count ---- # # The restored artoo_meta exposes the dataset-level attributes. get_meta(back)@dataset$records #> [1] 60 ``` # Read a dataset from CDISC Dataset-JSON NDJSON Read a newline-delimited CDISC Dataset-JSON v1.1 (`.ndjson`) file back to a data frame, restoring the full `artoo_meta` from its metadata line and realizing SAS date/datetime/time variables to R `Date` / `POSIXct` / [`hms::hms`](https://hms.tidyverse.org/reference/hms.html). Rows are parsed in bounded slabs, and `n_max` stops the line loop early, so a partial read of a huge file never parses the tail. A thin wrapper over [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) with `format = "ndjson"`. ## Usage ``` r read_ndjson(path, col_select = NULL, n_max = Inf, encoding = NULL) ``` ## Arguments - path: *Source `.ndjson` path.* `: required`. A gzip stream (`.ndjson.gz`) is inflated transparently. A file whose first line is not the Dataset-JSON metadata object aborts with `artoo_error_codec`. - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names. Columns return in file order (not the requested order) and the `artoo_meta` is filtered to match. Works on every format: parquet narrows columns natively, the rest filter after decode. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. xpt v8 bounds the disk read; the other formats cap after decode. - encoding: *Source charset of the file bytes.* ` | NULL`. `NULL` (default) reads UTF-8, as Dataset-JSON requires. Pass an IANA or SAS charset name (e.g. `"windows-1252"`) only to read a non-conformant file a producer wrote in that charset; each line is transcoded to UTF-8 on read, preserving the bounded `n_max` streaming. ## Value *A ``* carrying `artoo_meta` (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)). ## See also [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md) for the inverse; [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md) for the array-form file; [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: round-trip a conformed dataset through NDJSON ---- # # The variable labels, types, and keys survive the round-trip. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".ndjson") write_ndjson(adsl, path) back <- read_ndjson(path) identical(get_meta(back)@columns, get_meta(adsl)@columns) #> [1] TRUE # ---- Example 2: a bounded partial read of the first rows ---- # # n_max stops the line loop as soon as enough rows are in. head_rows <- read_ndjson(path, n_max = 5) get_meta(head_rows)@dataset$records #> [1] 5 ``` # Read a dataset from Apache Parquet Read an Apache Parquet (`.parquet`) file back to a data frame, restoring the `artoo_meta` from its `metadata_json` sidecar and realizing SAS date/datetime/time variables to R `Date` / `POSIXct` / [`hms::hms`](https://hms.tidyverse.org/reference/hms.html). A parquet written by another tool (with no artoo sidecar) reads back as a bare frame. A thin wrapper over [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) with `format = "parquet"`. Requires the lightweight `nanoparquet` package. ## Usage ``` r read_parquet(path, col_select = NULL, n_max = Inf, encoding = NULL) ``` ## Arguments - path: *Source `.parquet` path.* `: required`. - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names. Columns return in file order (not the requested order) and the `artoo_meta` is filtered to match. Works on every format: parquet narrows columns natively, the rest filter after decode. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. xpt v8 bounds the disk read; the other formats cap after decode. - encoding: *Source charset of the string columns.* ` | NULL`. `NULL` (default) reads the UTF-8 bytes parquet stores. Pass a charset name only to read a foreign file whose string columns hold that charset's bytes; they are transcoded to UTF-8 on read. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. ## Value *A ``* carrying `artoo_meta` when the file recorded it (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)); otherwise a plain data frame. ## See also [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md) for the inverse; [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: round-trip a conformed dataset through Parquet ---- # # The variable labels, types, and keys survive the round-trip. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".parquet") write_parquet(adsl, path) back <- read_parquet(path) get_meta(back)@columns$STUDYID$label #> [1] "Study Identifier" # ---- Example 2: the metadata names the dataset and row count ---- # # The restored artoo_meta exposes the dataset-level attributes. get_meta(back)@dataset$records #> [1] 60 ``` # Read a dataset from rds Read an R `.rds` file written by [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) (or any rds carrying a `metadata_json` attribute) back to a data frame with its `artoo_meta` restored. A thin wrapper over [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) with `format = "rds"`. ## Usage ``` r read_rds(path, col_select = NULL, n_max = Inf, encoding = NULL) ``` ## Arguments - path: *Source `.rds` path.* `: required`. - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names. Columns return in file order (not the requested order) and the `artoo_meta` is filtered to match. Works on every format: parquet narrows columns natively, the rest filter after decode. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. xpt v8 bounds the disk read; the other formats cap after decode. - encoding: *Source charset of the string columns.* ` | NULL`. `NULL` (default) returns the strings exactly as saved (faithful R round-trip). Pass a charset name only to transcode a foreign rds whose string columns hold that charset's bytes. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. ## Value *A ``* carrying `artoo_meta` when the file recorded it. An rds holding anything other than a data frame is a `artoo_error_codec`; use [`readRDS()`](https://rdrr.io/r/base/readRDS.html) for arbitrary objects. ## See also [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) for the inverse; [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) # ---- Example 1: read a dataset written by write_rds() ---- # # The restored frame carries the same metadata it was written with. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".rds") write_rds(adsl, path) back <- read_rds(path) get_meta(back)@dataset$records #> [1] 60 # ---- Example 2: a plain rds still reads as a data frame ---- # # An rds without artoo metadata reads back as an ordinary frame. bare <- tempfile(fileext = ".rds") saveRDS(cdisc_dm, bare) nrow(read_rds(bare)) #> [1] 60 ``` # Read a specification from JSON, Excel, or Define-XML Read a clinical-dataset specification into a validated `artoo_spec`, dispatching on the file extension: artoo's native JSON (the inverse of [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md)), a Pinnacle 21 (P21) Excel workbook, or a native Define-XML 2.0/2.1 document. The returned spec is the lingua franca the rest of artoo applies and serialises. ## Usage ``` r read_spec(path, datasets = NULL, on_duplicate = c("error", "first", "warn")) ``` ## Arguments - path: *The specification file to read.* `: required`. A `.json` (native) or `.xlsx` / `.xls` (P21) file. **Requirement:** reading a P21 workbook needs the `readxl` package. - datasets: *Read only these datasets.* ` | NULL`. `NULL` (default) reads the whole spec. Otherwise the spec is scoped to the named datasets before validation, so one broken sheet elsewhere in a workbook cannot block the dataset you are working on. An unknown name aborts listing what the file defines. - on_duplicate: *Policy for a variable defined more than once.* ``. A workbook row duplicated within one dataset makes the spec ambiguous; the finding is reported with its source location (sheet and row numbers for Excel). One of: - `"error"` (default) abort, naming each duplicate's rows. - `"first"` keep the first definition of each, dropping the rest with a message. - `"warn"` keep the first definition and warn (`artoo_warning_spec`). ## Value *A validated `artoo_spec`.* Inspect it with [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) / [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), check it with [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md), or persist it with [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md). ## Details **Three formats, one validator.** A `.json` file is read as artoo native JSON; a `.xlsx` / `.xls` file is read as a P21 workbook; a `.xml` file is read as Define-XML 2.x. Either way the result is built through [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md), so type canonicalisation and cross-slot integrity checks are identical regardless of source. **Define-XML ingestion** (needs the `xml2` package). ItemGroupDefs become datasets (keys derived from the ItemRef KeySequence), ItemRef + ItemDef pairs become variables, CodeLists become codelists (`def:ExtendedValue = "Yes"` marks an extended term), MethodDefs / CommentDefs / leaves become the supporting slots, and ValueListDefs land in the value-level slot with their where-clauses rendered as readable text. **Note:** an `ExternalCodeList` (MedDRA, ISO-3166) names a dictionary, not an enumerable membership list; it is dropped, and variables that referenced it carry no codelist. Define-XML v1.0 (the 2005 model) is refused with guidance. **P21 ingestion.** Sheets are located by a tolerant alias match (case-, space-, and spelling-variant insensitive). Datasets and Variables are required; Codelists and ValueLevel are optional (the latter becomes the spec's value-level slot). Every cell is read as text, then the dataset and codelist foreign keys are forward-filled to recover merged cells (which the Excel reader returns as `NA` on continuation rows). A key that cannot be resolved aborts with `artoo_error_spec` rather than being silently dropped. ## See also **Inverse:** [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) serialises a spec to native JSON. **Build / inspect:** [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md), [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md), [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## Examples ``` r # ---- Example 1: round-trip a spec through native JSON ---- # # write_spec() and read_spec() are inverses on the JSON path: the spec # that comes back is identical to the one written. spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) path <- tempfile(fileext = ".json") write_spec(spec, path) back <- read_spec(path) identical(back, spec) #> [1] TRUE # ---- Example 2: scope the read to one dataset ---- # # `datasets =` reads just the domain you are working on — validation is # scoped with it, so a problem elsewhere in the workbook cannot block # this dataset. dm_spec <- read_spec(path, datasets = "DM") spec_datasets(dm_spec) #> [1] "DM" head(spec_variables(dm_spec, "DM")[, c("variable", "label", "data_type")]) #> variable label data_type #> 1 STUDYID Study Identifier string #> 2 DOMAIN Domain Abbreviation string #> 3 USUBJID Unique Subject Identifier string #> 4 SUBJID Subject Identifier for the Study string #> 5 RFSTDTC Subject Reference Start Date/Time string #> 6 RFENDTC Subject Reference End Date/Time string ``` # Read a dataset from SAS XPORT Read a SAS Transport (`.xpt`) file (v5 or v8) back to a data frame, restoring the `artoo_meta` its NAMESTR records carry and realizing SAS date/datetime/time variables to R `Date` / `POSIXct` / [`hms::hms`](https://hms.tidyverse.org/reference/hms.html). The ingest end of the I/O layer; a thin wrapper over [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) with `format = "xpt"`. ## Usage ``` r read_xpt(path, encoding = NULL, col_select = NULL, n_max = Inf, member = NULL) ``` ## Arguments - path: *Source `.xpt` path.* `: required`. - encoding: *Force a source charset.* ` | NULL`. `NULL` (default) auto-detects (UTF-8 when every character value and label is valid UTF-8, else Windows-1252). IANA and SAS names both work. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. - col_select: *Variables to read.* ` | NULL`. `NULL` (default) reads every column; otherwise a vector of variable names (matching the names as stored, uppercase for v5). Columns return in file order, and the `artoo_meta` is filtered to match. **Note:** an unknown name is a `artoo_error_input`, never a silent drop. - n_max: *Maximum records to read.* `: default Inf`. Caps the row count; the returned `artoo_meta` reports the rows actually read. - member: *Which member of a multi-member transport file to read.* ` | NULL`. A transport file can hold several datasets; pass a member name (case-insensitive) or 1-based index to pick one. `NULL` (default) reads a single-member file directly and aborts on a multi-member file, pointing at [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md). **Tip:** `xpt_members(path)` lists what a file holds before you choose. ## Value *A ``* carrying `artoo_meta` (read it with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md)). ## Details The character encoding is auto-detected (UTF-8 if every character value is valid UTF-8, else Windows-1252) and recorded on the returned `artoo_meta`, so a later [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) reproduces it; pass `encoding` to override. XPORT cannot record its own encoding, so this detection is a heuristic. See [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) for what XPORT can and cannot preserve. ## See also [`xpt_members()`](https://vthanik.github.io/artoo/reference/xpt_members.md) to list a file's members; [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) for the inverse; [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) # ---- Example 1: round-trip a conformed dataset through xpt ---- # # Write ADSL, read it back; the variable labels and lengths survive. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".xpt") write_xpt(adsl, path) back <- read_xpt(path) get_meta(back)@columns$STUDYID$label #> [1] "Study Identifier" # ---- Example 2: pick one member of a multi-member transport file ---- # # Build a two-member file by concatenating two single-member files (every # member section is 80-byte padded), then read one dataset out of it. dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. p_dm <- tempfile(fileext = ".xpt") write_xpt(dm, p_dm) #> Warning: Widened 1 column past the declared spec length: "STUDYID (7 -> 12)". #> ℹ Values need more bytes than the spec length; data was kept whole. #> ℹ Update the spec length, or shorten the data, so the file matches its #> declared metadata. multi <- tempfile(fileext = ".xpt") writeBin( c( readBin(path, "raw", file.size(path)), readBin(p_dm, "raw", file.size(p_dm))[-(1:240)] ), multi ) xpt_members(multi)$name #> [1] "ADSL" "DM" nrow(read_xpt(multi, member = "DM")) #> [1] 60 ``` # Repair a spec from its conformance findings Take the `integer_fraction` and `integer_overflow` findings a [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) or [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) run reports and return a new spec with every offending variable retyped to `"float"`, so a frame that the original spec would refuse to coerce now conforms. This closes the loop on the spec-side fix: inspect the findings, then apply them all at once instead of editing the source workbook variable by variable. Persist the result with [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md). ## Usage ``` r repair_spec(spec, findings) ``` ## Arguments - spec: *The specification to repair.* `: required`. - findings: *A findings data frame.* `: required`. The result of [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) or [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md); must carry the `check`, `dataset`, and `variable` columns. ## Value *A new ``* with the flagged variables retyped to `"float"`, or `spec` unchanged when there is nothing to repair. The input is never mutated. ## Details **Scope.** Only the two lossy-integer findings are repaired (`integer_fraction`, `integer_overflow`) — both mean "the spec says `integer` but the data is not", and `"float"` is the loss-free fix. Other findings are ignored; this is not a general spec rewriter. When no repairable finding is present the spec is returned unchanged, with a note. **Built on [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md).** Each `(dataset, variable)` pair is applied through the same validated override [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) uses, so the result is a fully re-validated `artoo_spec`, never a hand-edited internal. ## See also **Primitive:** [`set_type()`](https://vthanik.github.io/artoo/reference/set_type.md) to retype a chosen variable directly. **Findings:** [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) for one dataset, [`check_study()`](https://vthanik.github.io/artoo/reference/check_study.md) across a study. **Persist:** [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md). ## Examples ``` r # ---- Example 1: auto-repair an integer/fractional mismatch ---- # # adam_spec types ADSL.AGE as integer. Give it fractional ages and # check_spec() raises an integer_fraction error; repair_spec() flips AGE # (and only AGE) to float, and the corrected spec then applies cleanly. dat <- cdisc_adsl dat$AGE <- dat$AGE + 0.5 findings <- check_spec(dat, adam_spec, "ADSL") fixed <- repair_spec(adam_spec, findings) spec_variables(fixed, "ADSL")$data_type[ spec_variables(fixed, "ADSL")$variable == "AGE" ] #> [1] "float" # ---- Example 2: nothing to repair is a no-op ---- # # The bundled data conforms, so its findings carry no integer_fraction or # integer_overflow rows and the spec is returned unchanged. clean <- check_spec(cdisc_adsl, adam_spec, "ADSL") identical(repair_spec(adam_spec, clean), adam_spec) #> No "integer_fraction" or "integer_overflow" findings to repair. #> ℹ The spec is returned unchanged. #> [1] TRUE ``` # Attach metadata to a dataset Stamp a `artoo_meta` onto a data frame as a single Dataset-JSON string in its `metadata_json` attribute. Every `write_*()` codec reads that string back with [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) and embeds it verbatim, so the metadata survives the trip to any format. Use it to attach metadata to a bare frame before a write, or to re-stamp after a tidyverse verb has dropped attributes. ## Usage ``` r set_meta(x, meta) ``` ## Arguments - x: *The data frame to stamp.* `: required`. - meta: *The metadata to attach.* `: required`. Usually from [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) or built by [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md). ## Value *The data frame `x`*, with its `metadata_json` attribute set. Pass it on to a `write_*()` codec or back through [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md). ## See also [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) for the read half; [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) which stamps it. ## Examples ``` r # ---- Example 1: re-stamp metadata a dplyr verb would drop ---- # # Conform a dataset, capture its metadata, then re-attach after an # attribute-dropping transform so the write stays lossless. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) adsl <- apply_spec(cdisc_adsl, spec, "ADSL") meta <- get_meta(adsl) trimmed <- head(as.data.frame(adsl), 5) attr(trimmed, "metadata_json") <- NULL set_meta(trimmed, meta) #> STUDYID USUBJID SUBJID SITEID SITEGR1 ARM #> 1 CDISCPILOT01 01-701-1015 1015 701 701 Placebo #> 2 CDISCPILOT01 01-701-1023 1023 701 701 Placebo #> 3 CDISCPILOT01 01-701-1028 1028 701 701 Xanomeline High Dose #> 4 CDISCPILOT01 01-701-1033 1033 701 701 Xanomeline Low Dose #> 5 CDISCPILOT01 01-701-1034 1034 701 701 Xanomeline High Dose #> TRT01P TRT01PN TRT01A TRT01AN TRTSDT #> 1 Placebo 0 Placebo 0 2014-01-02 #> 2 Placebo 0 Placebo 0 2012-08-05 #> 3 Xanomeline High Dose 81 Xanomeline High Dose 81 2013-07-19 #> 4 Xanomeline Low Dose 54 Xanomeline Low Dose 54 2014-03-18 #> 5 Xanomeline High Dose 81 Xanomeline High Dose 81 2014-07-01 #> TRTEDT TRTDUR AVGDD CUMDOSE AGE AGEGR1 AGEGR1N AGEU RACE RACEN #> 1 2014-07-02 182 0.0 0 63 <65 1 YEARS WHITE 1 #> 2 2012-09-01 28 0.0 0 64 <65 1 YEARS WHITE 1 #> 3 2014-01-14 180 77.7 13986 71 65-80 2 YEARS WHITE 1 #> 4 2014-03-31 14 54.0 756 74 65-80 2 YEARS WHITE 1 #> 5 2014-12-30 183 76.9 14067 77 65-80 2 YEARS WHITE 1 #> SEX ETHNIC SAFFL ITTFL EFFFL COMP8FL COMP16FL #> 1 F HISPANIC OR LATINO Y Y Y Y Y #> 2 M HISPANIC OR LATINO Y Y Y N N #> 3 M NOT HISPANIC OR LATINO Y Y Y Y Y #> 4 M NOT HISPANIC OR LATINO Y Y Y N N #> 5 F NOT HISPANIC OR LATINO Y Y Y Y Y #> COMP24FL DISCONFL DSRAEFL DTHFL BMIBL BMIBLGR1 HEIGHTBL WEIGHTBL #> 1 Y 25.1 25-<30 147.3 54.4 #> 2 N Y Y 30.4 >=30 162.6 80.3 #> 3 Y 31.4 >=30 177.8 99.3 #> 4 N Y 28.8 25-<30 175.3 88.5 #> 5 Y 26.1 25-<30 154.9 62.6 #> EDUCLVL DISONSDT DURDIS DURDSGR1 VISIT1DT RFSTDTC RFENDTC #> 1 16 2010-04-30 43.9 >=12 2013-12-26 2014-01-02 2014-07-02 #> 2 14 2006-03-11 76.4 >=12 2012-07-22 2012-08-05 2012-09-02 #> 3 16 2009-12-16 42.8 >=12 2013-07-11 2013-07-19 2014-01-14 #> 4 12 2009-08-02 55.3 >=12 2014-03-10 2014-03-18 2014-04-14 #> 5 9 2011-09-29 32.9 >=12 2014-06-24 2014-07-01 2014-12-30 #> VISNUMEN RFENDT DCDECOD DCREASCD #> 1 12 2014-07-02 COMPLETED Completed #> 2 5 2012-09-02 ADVERSE EVENT Adverse Event #> 3 12 2014-01-14 COMPLETED Completed #> 4 5 2014-04-14 STUDY TERMINATED BY SPONSOR Sponsor Decision #> 5 12 2014-12-30 COMPLETED Completed #> MMSETOT #> 1 23 #> 2 23 #> 3 23 #> 4 23 #> 5 21 # ---- Example 2: borrow metadata from a conformed dataset ---- # # A writer with a raw frame can lift metadata off a conformed dataset and # stamp it onto the bare frame (DM is SDTM, so its spec is sdtm-shaped). sdtm <- artoo_spec( cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists ) meta_dm <- get_meta(apply_spec(cdisc_dm, sdtm, "DM")) dm <- set_meta(cdisc_dm, meta_dm) is_artoo_meta(get_meta(dm)) #> [1] TRUE ``` # Override a variable's dataType in a spec Return a new `artoo_spec` with one or more variables retyped. This is the supported, in-R way to correct a spec when the data disagrees with its declared `dataType` (e.g. a variable typed `integer` whose extract holds fractional values): fix it here rather than editing the source workbook, then drive [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) with the corrected spec. The spec is immutable, so the original is never changed. ## Usage ``` r set_type(spec, dataset, ...) ``` ## Arguments - spec: *The specification to amend.* `: required`. - dataset: *The dataset whose variables to retype.* `: required`. Must name a dataset in `spec`. - ...: *Named `variable = type` pairs.* Each name is a variable in `dataset`; each value is a CDISC `dataType` (`"string"`, `"integer"`, `"decimal"`, `"float"`, `"double"`, `"boolean"`, `"date"`, `"datetime"`, `"time"`, `"URI"`) or a recognised spelling of one. At least one pair is required and every argument must be named. **Tip:** to undo an `integer` dataType that the data does not satisfy, set `"float"` (IEEE double) or `"decimal"` (exact, exchanged as text). ## Value *A new ``* with the named variables retyped, ready for [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) or [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md). The input `spec` is unchanged. ## Details **Per-dataset scope.** A type is set only on the named `dataset`'s row. A variable that appears in several datasets keeps its other rows' types; call `set_type()` once per dataset to change them all. Spec-wide consequences (a variable typed inconsistently across datasets) are a [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) concern, not a construction error here. **Canonicalised, then validated.** Each supplied type is mapped through the closed CDISC `dataType` vocabulary, so `"Float"`, `"decimal"`, and `"text"` all resolve; an unrecognised token aborts with `artoo_error_type`. The rebuilt spec is re-validated, so an override that would break the spec aborts with `artoo_error_spec`. ## See also **Auto-repair:** [`repair_spec()`](https://vthanik.github.io/artoo/reference/repair_spec.md) to apply every `integer_fraction` / `integer_overflow` fix from a findings frame at once. **Workflow:** [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) to conform with the corrected spec; [`write_spec()`](https://vthanik.github.io/artoo/reference/write_spec.md) to persist it; [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md) to find the mismatches. ## Examples ``` r # ---- Example 1: retype one variable the data disagrees with ---- # # The bundled adam_spec types ADSL.AGE as integer. If an extract stored it # with fractional values, retype it to float so apply_spec() coerces # without loss. set_type() returns a new spec; the original is untouched. fixed <- set_type(adam_spec, "ADSL", AGE = "float") v <- spec_variables(fixed, "ADSL") v[v$variable == "AGE", c("variable", "data_type")] #> variable data_type #> 16 AGE float # ---- Example 2: retype several at once, original left intact ---- # # Pass any number of variable = type pairs; canonical dataTypes and common # spellings both resolve. The source spec is immutable, so adam_spec still # reports AGE as its original type. patched <- set_type(adam_spec, "ADSL", AGE = "decimal", TRTSDT = "date") spec_variables(adam_spec, "ADSL")$data_type[ spec_variables(adam_spec, "ADSL")$variable == "AGE" ] #> [1] "integer" ``` # Codelist terms Return the controlled-terminology terms and decodes a spec carries: one codelist's terms when `codelist_id` names it, or the full `codelists` slot when `codelist_id` is `NULL`. Use it to inspect the values a coded variable is allowed to take before applying the spec. Mirrors the [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md) filter pattern. ## Usage ``` r spec_codelists(spec, codelist_id = NULL) ``` ## Arguments - spec: *The specification to read.* `: required`. - codelist_id: *The codelist to return.* ` | NULL`. When `NULL` (default) the whole codelists table is returned. **Restriction:** a non-`NULL` id must name a codelist present in the spec's `codelists` slot; an unknown id aborts with `artoo_error_input`. ## Value *A data frame of codelist terms*, one row per term: every term when `codelist_id` is `NULL`, else the named codelist's terms. Columns: - `codelist_id` — the codelist identifier variables reference. - `term` — the submission value (what conformed data carries). - `decode` — the human-readable decoded value. - `order` — display order within the codelist. - `extended` — `TRUE` marks an extensible codelist (sponsor terms allowed; non-members downgrade to notes in [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md)). - `comment_id` — reference into the comments slot. ## See also [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md) for which variables reference a codelist. ## Examples ``` r # ---- Example 1: the terms behind a coded variable ---- # # SEX is coded against C66731; spec_codelists() returns the terms and their # decodes that apply_spec() will enforce or decode. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) spec_codelists(spec, "C66731") #> codelist_id term decode order extended #> 1 C66731 F Female 1 NA #> 2 C66731 M Male 2 NA #> 3 C66731 U Unknown 3 NA #> 4 C66731 UNDIFFERENTIATED Undifferentiated 4 NA #> comment_id #> 1 #> 2 #> 3 #> 4 # ---- Example 2: the whole codelists table ---- # # Called with no id, it returns every term across every codelist. head(spec_codelists(spec)) #> codelist_id term decode order extended #> 1 C66731 F Female 1 NA #> 2 C66731 M Male 2 NA #> 3 C66731 U Unknown 3 NA #> 4 C66731 UNDIFFERENTIATED Undifferentiated 4 NA #> comment_id #> 1 #> 2 #> 3 #> 4 ``` # Comment definitions in a spec Return the comment definitions a specification carries. Datasets, variables, value-level rows, and codelists reference these by `comment_id`; [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) checks the references resolve and each referenced comment has a body. ## Usage ``` r spec_comments(spec) ``` ## Arguments - spec: *The specification to read.* `: required`. ## Value *A data frame of comment metadata*, one row per comment, with all four columns: `comment_id`, `description`, `document_id`, `pages`. Empty when the spec defines no comments. ## See also [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md), [`spec_documents()`](https://vthanik.github.io/artoo/reference/spec_documents.md), [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## Examples ``` r # ---- Example 1: the comments a spec defines ---- # # Build a spec with one comment and read it back. spec <- artoo_spec( data.frame(dataset = "ADSL"), data.frame(dataset = "ADSL", variable = "AGE", data_type = "integer"), comments = data.frame( comment_id = "C.AGE", description = "Age in years at informed consent.", stringsAsFactors = FALSE ) ) spec_comments(spec) #> comment_id description document_id pages #> 1 C.AGE Age in years at informed consent. ``` # Dataset names in a spec List the datasets a specification defines. The result is the set of names you pass as the `dataset` argument to the other accessors and to [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md). ## Usage ``` r spec_datasets(spec) ``` ## Arguments - spec: *The specification to read.* `: required`. ## Value *A character vector of dataset names*, de-duplicated and with `NA`s dropped. Empty when the spec has no datasets. ## See also [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md) for one dataset's variables; [`spec_keys()`](https://vthanik.github.io/artoo/reference/spec_keys.md) for its sort keys. ## Examples ``` r # ---- Example 1: the datasets the pilot ADaM spec defines ---- # # Build the spec from the bundled CDISC-pilot tables and list its # datasets — the names you pass to the other accessors. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) spec_datasets(spec) #> [1] "ADSL" ``` # Document references in a spec Return the document references a specification carries. Methods and comments point to these by `document_id`. ## Usage ``` r spec_documents(spec) ``` ## Arguments - spec: *The specification to read.* `: required`. ## Value *A data frame of document metadata* (`document_id`, `title`, `href`), one row per document. Empty when the spec defines none. ## See also [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md), [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md), [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## Examples ``` r # ---- Example 1: the documents a spec defines ---- # # Build a spec with one document reference and read it back. spec <- artoo_spec( data.frame(dataset = "ADSL"), data.frame(dataset = "ADSL", variable = "AGE", data_type = "integer"), documents = data.frame( document_id = "SAP", title = "Statistical Analysis Plan", stringsAsFactors = FALSE ) ) spec_documents(spec) #> document_id title href #> 1 SAP Statistical Analysis Plan ``` # Sort keys for a dataset Parse a dataset's sort keys into a character vector of variable names. These keys drive the sort step of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md) and the `keySequence` written to each output format. ## Usage ``` r spec_keys(spec, dataset) ``` ## Arguments - spec: *The specification to read.* `: required`. - dataset: *The dataset whose keys to parse.* `: required`. **Restriction:** must name a dataset in the spec. ## Value *A character vector of key variable names*, split from the dataset's `keys` cell (whitespace- or comma-separated). Empty when no keys are declared. ## See also [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) for the dataset names; [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md) for the variables a key must reference. ## Examples ``` r # ---- Example 1: parse a dataset's sort keys ---- # # Declare DM's keys, then read them back as the ordered vector apply_spec() # sorts by. (STUDYID and USUBJID are real DM variables in the demo data.) ds <- cdisc_sdtm_datasets ds$keys[ds$dataset == "DM"] <- "STUDYID USUBJID" spec <- artoo_spec(ds, cdisc_sdtm_variables, codelists = cdisc_codelists) spec_keys(spec, "DM") #> [1] "STUDYID" "USUBJID" ``` # Derivation methods in a spec Return the method definitions a specification carries. Variables and value-level rows reference these by `method_id`; [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md) checks that every reference resolves and that each referenced method is complete (has a description). ## Usage ``` r spec_methods(spec) ``` ## Arguments - spec: *The specification to read.* `: required`. ## Value *A data frame of method metadata*, one row per method, with all eight columns: `method_id`, `description`, `name`, `type`, `expression_context`, `expression_code`, `document_id`, `pages`. Empty when the spec defines no methods. ## See also [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md), [`spec_documents()`](https://vthanik.github.io/artoo/reference/spec_documents.md), [`validate_spec()`](https://vthanik.github.io/artoo/reference/validate_spec.md). ## Examples ``` r # ---- Example 1: the methods a spec defines ---- # # Build a spec with one derivation method and read it back. spec <- artoo_spec( data.frame(dataset = "ADSL"), data.frame(dataset = "ADSL", variable = "AGEGR1", data_type = "string"), methods = data.frame( method_id = "MT.AGEGR1", description = "Age group from AGE.", stringsAsFactors = FALSE ) ) spec_methods(spec) #> method_id description name type expression_context #> 1 MT.AGEGR1 Age group from AGE. #> expression_code document_id pages #> 1 ``` # The CDISC standard a spec implements Return the one CDISC standard the specification carries (e.g. `"ADaMIG 1.1"`, `"SDTMIG 3.2"`). A `artoo_spec` is single-standard by construction — [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) aborts when its sources mix standards — so this is always a scalar; `NA` when no source named one. ## Usage ``` r spec_standard(spec) ``` ## Arguments - spec: *The specification to read.* `: required`. ## Value *A ``*: the standard, or `NA` when unspecified. ## See also [`spec_study()`](https://vthanik.github.io/artoo/reference/spec_study.md) for the rest of the study-level metadata; [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) for how the standard is resolved. ## Examples ``` r # ---- Example 1: the standard set at construction ---- # # Pass the standard explicitly (or let it resolve from a P21 workbook's # Standard column / a Define-XML study block) and read it back. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists, standard = "ADaMIG 1.1" ) spec_standard(spec) #> [1] "ADaMIG 1.1" # ---- Example 2: unspecified resolves to NA ---- # # A spec built without any standard source carries NA. bare <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) spec_standard(bare) #> [1] "ADaMIG 1.1" ``` # Study-level metadata Return the study-level metadata row, or a single field from it. Holds the canonical study fields (`study_name`, `study_description`, `protocol_name` — every source spelling is canonicalised by [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md)) plus any other study-scoped fields a source provides (the CDISC standard lives on its own property — see [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md)). ## Usage ``` r spec_study(spec, field = NULL) ``` ## Arguments - spec: *The specification to read.* `: required`. - field: *Return one field instead of the row.* ` | NULL`. When `NULL` (default) the whole study data frame is returned. **Restriction:** a non-`NULL` `field` must be a column of the study table; an unknown field aborts with `artoo_error_input`. ## Value *The study data frame* (one row), or the value of one `field`. The canonical fields are `study_name`, `study_description`, and `protocol_name` — every source spelling is canonicalised to these by [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) — plus any other field the source carried verbatim (e.g. `define_version` from a Define-XML read). ## See also [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) for the datasets the study scopes; [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md) for the spec's CDISC standard. ## Examples ``` r # ---- Example 1: the whole study row, then one field ---- # # spec_study() with no field returns the study-level table; pass a field # name to pull a single value such as the study name. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists, study = data.frame(study_name = "CDISCPILOT01") ) spec_study(spec) #> study_name #> 1 CDISCPILOT01 spec_study(spec, "study_name") #> [1] "CDISCPILOT01" ``` # Variables in a spec Return the variable-metadata table for one dataset, or for the whole spec. Each row carries the variable's CDISC `data_type`, label, length, display format, key sequence, and codelist reference. ## Usage ``` r spec_variables(spec, dataset = NULL) ``` ## Arguments - spec: *The specification to read.* `: required`. - dataset: *Restrict to one dataset.* ` | NULL`. When `NULL` (default) every dataset's variables are returned; otherwise only the named dataset's rows. **Restriction:** a non-`NULL` `dataset` must name a dataset in the spec (see [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md)); an unknown name aborts with `artoo_error_input`. ## Value *A data frame of variable metadata*, one row per variable, with 22 columns (absent ones are filled with typed `NA` at construction): - `dataset`, `variable` — the identifying pair (unique within a spec). - `itemoid` — the Define-XML / Dataset-JSON item OID, when recorded. - `label` — the variable label (\<= 40 bytes for XPORT v5). - `data_type` — canonical CDISC dataType (`string`, `integer`, `decimal`, `float`, `double`, `boolean`, `date`, `datetime`, `time`, `URI`). - `target_data_type` — `integer`/`decimal` when a temporal variable stores as a SAS-epoch numeric; `NA` means ISO 8601 text (`--DTC`). - `length` — declared storage length (bytes for character). - `display_format`, `informat` — SAS format / informat strings. - `key_sequence` — 1-based position in the dataset sort key. - `order` — column position in the dataset. - `codelist_id`, `method_id`, `comment_id` — references into the codelists / methods / comments slots. - `mandatory` — logical obligation flag (`NA` is treated as mandatory by [`check_spec()`](https://vthanik.github.io/artoo/reference/check_spec.md)). - `significant_digits` — for `decimal` variables. - `origin`, `source`, `predecessor`, `assigned_value`, `pages`, `role` — Define-XML provenance fields, carried as-is. Filter or arrange it with ordinary base / `dplyr` verbs. ## See also [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md) for the dataset names; [`spec_codelists()`](https://vthanik.github.io/artoo/reference/spec_codelists.md) for a variable's controlled terminology. ## Examples ``` r spec <- artoo_spec(cdisc_sdtm_datasets, cdisc_sdtm_variables, codelists = cdisc_codelists) # ---- Example 1: one dataset's variables ---- # # Pass a dataset name to get just that domain's variables, already # canonicalised to CDISC dataTypes. head(spec_variables(spec, "DM")[, c("variable", "label", "data_type")]) #> variable label data_type #> 1 STUDYID Study Identifier string #> 2 DOMAIN Domain Abbreviation string #> 3 USUBJID Unique Subject Identifier string #> 4 SUBJID Subject Identifier for the Study string #> 5 RFSTDTC Subject Reference Start Date/Time string #> 6 RFENDTC Subject Reference End Date/Time string # ---- Example 2: every variable across the spec ---- # # Omit `dataset` to get the full table, e.g. to count variables per domain. table(spec_variables(spec)$dataset) #> #> DM #> 25 ``` # Re-align metadata with a transformed data frame Re-attach and reconcile a `artoo_meta` after a transformation that dropped or reshaped it: the metadata's columns are narrowed and reordered to the frame's current columns, the record count is refreshed, the keys are recomputed, and a column the metadata does not describe gets an entry synthesized from its class and attributes. The one-liner to run after a dplyr (or base) pipeline, before handing the frame to a `write_*()` codec. ## Usage ``` r sync_meta(x, meta = NULL) ``` ## Arguments - x: *The transformed data frame.* `: required`. - meta: *The metadata to reconcile against.* ` | NULL`. `NULL` (default) uses the frame's own `metadata_json` attribute. **Requirement:** when the transform dropped the attribute (base `[` subsetting does), capture [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md) before the pipeline and pass it here; a bare frame with no `meta` aborts with `artoo_error_input`. ## Value *A ``*: `x` re-stamped with the reconciled `artoo_meta`. Hand it to any `write_*()` codec. ## Details **Why it exists.** Base row subsetting (`x[i, ]`) drops the frame's `metadata_json` attribute, and many tidyverse verbs rebuild the frame. `sync_meta()` takes the last-known metadata (the frame's own attribute when it survived, or an explicit `meta`) and makes it agree with the data again, so the round trip stays lossless without hand-editing. ## See also **Read / attach:** [`get_meta()`](https://vthanik.github.io/artoo/reference/get_meta.md), [`set_meta()`](https://vthanik.github.io/artoo/reference/set_meta.md). **Produce conformed frames:** [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md). ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: re-attach after an attribute-dropping subset ---- # # Base subsetting drops the metadata; capture it first, transform, then # sync. The metadata narrows to the kept columns and the new row count. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") meta <- get_meta(adsl) elderly <- adsl[adsl$AGE > 65, c("STUDYID", "USUBJID", "AGE")] synced <- sync_meta(elderly, meta) get_meta(synced)@dataset$records #> [1] 46 # ---- Example 2: a derived column gains a synthesized entry ---- # # A new column the metadata does not describe is profiled from its class, # so the frame still writes losslessly. adsl$AGEGR9 <- ifelse(adsl$AGE > 65, ">65", "<=65") synced2 <- sync_meta(adsl) #> Synthesized metadata for 1 new column: `AGEGR9`. #> ℹ Edit the spec (or the meta) if the inferred types need refining. get_meta(synced2)@columns$AGEGR9$dataType #> [1] "string" ``` # Validate a specification for submission-readiness Run artoo's bundled, self-contained checks over a `artoo_spec`, **scoped to the dataset(s) you are working on**, and return a `artoo_check` that prints a sectioned report. Every finding is keyed to an open rule in the shipped catalog (see `spec_rules.json`); the result object keeps the findings as a plain data frame in `@findings` for programmatic use. ## Usage ``` r validate_spec( spec, data = NULL, dataset = NULL, on_error = c("off", "warn", "abort") ) ``` ## Arguments - spec: *The specification to validate.* `: required`. - data: *Optional input data for controlled-terminology checks.* ` | named list | NULL`. When supplied, data values are cross-checked against the spec codelists. A single data frame requires a length-1 `dataset`; pass a named list to validate several at once. - dataset: *Restrict to one or more datasets.* ` | NULL`. `NULL` (default) validates every dataset. **Restriction:** each name must be a dataset in the spec. - on_error: *What to do with an error-severity finding.* ``. One of: - `"off"` (default) collect and return every finding; never signal. - `"warn"` additionally `cli_warn` (`artoo_warning_validation`) with the error count. - `"abort"` additionally abort with `artoo_error_validation`. All findings are collected and returned in every case. ## Value *A `artoo_check` object.* Its `@findings` data frame has columns `check`, `dimension`, `severity`, `dataset`, `variable`, `message`. Print it for the sectioned report. ## Details **Dataset-scoped.** A spec workbook carries many datasets. Pass `dataset` to validate only the one(s) you are working on — the methods, comments, and codelists those datasets reference are checked for completeness, but unrelated datasets are not. `dataset = NULL` validates the whole spec. **Collect, do not stop.** Every finding is collected and returned; `validate_spec()` does not abort on an error-severity finding unless `on_error = "abort"`. ## See also [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) to build a spec; [`spec_methods()`](https://vthanik.github.io/artoo/reference/spec_methods.md) / [`spec_comments()`](https://vthanik.github.io/artoo/reference/spec_comments.md) for the metadata checked. ## Examples ``` r # ---- Example 1: validate one dataset ---- # # Build a spec from the bundled ADaM tables and validate it; the # result prints a sectioned report and keeps the findings table. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) chk <- validate_spec(spec, dataset = "ADSL") chk@findings #> check dimension severity dataset variable #> 1 study_name_present study warning #> message #> 1 No study-level metadata; the study name is unknown. # ---- Example 2: gate on errors with on_error = "abort" ---- # # Point a key at a missing variable, then validate with on_error = "abort" # and catch the resulting error. bad_ds <- cdisc_sdtm_datasets bad_ds$keys <- "NOTAVAR" bad <- artoo_spec(bad_ds, cdisc_sdtm_variables, codelists = cdisc_codelists) tryCatch( validate_spec(bad, dataset = "DM", on_error = "abort"), artoo_error_validation = function(e) conditionMessage(e)[1] ) #> [1] "\033[1m\033[22mSpec is not submission-ready, 1 error-severity finding.\n\033[31m✖\033[39m Dataset 'DM' keys reference variables not in the spec: NOTAVAR.\n\033[36mℹ\033[39m Inspect every finding in the returned artoo_check." ``` # Write a dataset to any supported format Serialize a data frame to a clinical file format, preserving its `artoo_meta` losslessly. The codec is chosen from the file extension (or an explicit `format`), so one call covers xpt, Dataset-JSON, Parquet, and rds. This is the emit end of the artoo workflow; the per-format wrappers like [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) are thin sugar over it. ## Usage ``` r write_dataset(x, path, format = NULL, ...) ``` ## Arguments - x: *The dataset to write.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), carrying `artoo_meta`. - path: *Destination file path.* `: required`. Its extension selects the codec unless `format` is given. - format: *Force a codec instead of inferring from the extension.* ` | NULL`. One of the registered formats (see [`artoo_formats()`](https://vthanik.github.io/artoo/reference/artoo_formats.md)). - ...: *Codec-specific arguments* passed through to the encoder (see the per-format wrappers, e.g. [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md), for what each codec accepts). An argument the codec does not know is an error, never silently ignored. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. Called for the side effect of writing `path`. ## See also [`read_dataset()`](https://vthanik.github.io/artoo/reference/read_dataset.md) for the inverse; [`write_rds()`](https://vthanik.github.io/artoo/reference/write_rds.md) for the per-format wrapper; [`artoo_formats()`](https://vthanik.github.io/artoo/reference/artoo_formats.md) for what is available. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: write a conformed dataset, inferring rds from the path ---- # # apply_spec() attaches the metadata; write_dataset() carries it into the # file so a later read is lossless. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".rds") write_dataset(adsl, path) # ---- Example 2: force the format for an unconventional extension ---- # # When the extension does not name the format, pass it explicitly. alt <- tempfile(fileext = ".data") write_dataset(adsl, alt, format = "rds") ``` # Write a dataset to CDISC Dataset-JSON Serialize a data frame to a CDISC Dataset-JSON v1.1 (`.json`) file, Dataset-JSON being the native home of the `artoo_meta` shape: the file is the metadata block plus a flat `rows` array. The emit end of the artoo workflow (spec -\> apply_spec -\> write_json); a thin wrapper over [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) with `format = "json"`. ## Usage ``` r write_json( x, path, on_invalid = c("error", "translit", "fold", "replace", "ignore"), created = NULL, strict = FALSE ) ``` ## Arguments - x: *The dataset to write.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), carrying `artoo_meta`. - path: *Destination `.json` path.* `: required`. - on_invalid: *Policy for values that are not valid UTF-8.* `: default "error"`. One of `"error"` (abort with `artoo_error_codec`, naming the offenders with their invalid bytes hex-escaped), `"replace"` (substitute `?` and warn with `artoo_warning_encoding`), `"ignore"` (drop the invalid bytes), or `"translit"` / `"fold"` (like `"error"`; a byte-level invalidity has no character fold, the options exist so one policy value can thread a whole multi-format pipeline). The same policy vocabulary as [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md); text correctly read through artoo is always valid UTF-8, so this only fires on bytes that entered the frame through a mis-declared source encoding. - created: *Creation timestamp.* ` | NULL`. `NULL` (default) stamps the current time into `datasetJSONCreationDateTime`; freeze it for byte-stable output. - strict: *Suppress the `_artoo` extension block.* `: default FALSE`. By default the file carries a single namespaced `_artoo` object when (and only when) there is content strict CDISC cannot express: SAS special-missing tags (`.A`-`.Z`, `._`), the recorded source encoding, and informats. Data values stay plain `null`s either way, so a foreign reader degrades gracefully. **Note:** `strict = TRUE` writes a pure closed-vocabulary file and warns (`artoo_warning_codec`) naming exactly what was dropped; those attributes will not survive a read-back. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. ## Details **Full metadata, no loss.** Unlike `.xpt`, a `.json` file records the complete `artoo_meta`: keySequence, codelist, origin, targetDataType, and significantDigits all survive. Dates, datetimes, and times are exchanged as ISO 8601 strings, or as SAS-epoch numbers when their `targetDataType` is `"integer"` (the ADaM numeric-date convention); `decimal` rides as a string so exact precision is preserved. The file is always UTF-8 (RFC 8259 / CDISC v1.1). `NaN` and infinite values are not valid CDISC numerics and abort the write. **Streaming write, whole-file read.** The writer streams the `rows` array in bounded slabs (a `.json.gz` path gzips the stream transparently), but [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md) must parse the whole array at once. For multi-million-row datasets prefer the NDJSON variant ([`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md) / [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md)), which bounds memory in both directions. ## See also [`read_json()`](https://vthanik.github.io/artoo/reference/read_json.md) for the inverse; [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the generic dispatcher. ## Examples ``` r # ---- Example 1: write a conformed dataset as Dataset-JSON ---- # # apply_spec() attaches the metadata; write_json() serializes the full # itemGroup plus the data rows. adsl <- apply_spec(cdisc_adsl, adam_spec, "ADSL", conformance = "off") #> 6 variables the spec declares are absent from the data (not added): #> `TRTDURD`, `DISONDT`, `EOSSTT`, `DCSREAS`, `EOSDISP`, and `MMS1TSBL`. path <- tempfile(fileext = ".json") write_json(adsl, path) # ---- Example 2: a frozen timestamp for reproducible bytes ---- # # Fixing `created` makes two writes byte-identical; the columns() pane on # the written file shows the full metadata the file carries (DM is SDTM, # so it conforms against the bundled sdtm_spec). dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. path2 <- tempfile(fileext = ".json") write_json(dm, path2, created = as.POSIXct("2020-01-01", tz = "UTC")) columns(path2) #> DM -- 25 variables, 60 obs #> # Variable Type Len Format Label Key #> 1 STUDYID Char 7 Study Identifier 1 #> 2 DOMAIN Char 2 Domain Abbreviation #> 3 USUBJID Char 14 Unique Subject Identifier 2 #> 4 SUBJID Char 6 Subject Identifier for the Study #> 5 RFSTDTC Char 10 Subject Reference Start Date/Time #> 6 RFENDTC Char 10 Subject Reference End Date/Time #> 7 SITEID Char 3 Study Site Identifier #> 8 AGE Num Age #> 9 AGEU Char 5 Age Units #> 10 SEX Char 16 Sex #> 11 RACE Char 41 Race #> 12 ETHNIC Char 22 Ethnicity #> 13 ARMCD Char 8 Planned Arm Code #> 14 ARM Char 20 Description of Planned Arm #> 15 COUNTRY Char 3 Country #> 16 RFXSTDTC Char 10 #> 17 RFXENDTC Char 10 #> 18 RFICDTC Char 1 #> 19 RFPENDTC Char 16 #> 20 DTHDTC Char 10 #> 21 DTHFL Char 1 #> 22 ACTARMCD Char 8 #> 23 ACTARM Char 20 #> 24 DMDTC Char 10 #> 25 DMDY Num ``` # Write a dataset to CDISC Dataset-JSON NDJSON Serialize a data frame to the newline-delimited variant of CDISC Dataset-JSON v1.1 (`.ndjson`): line 1 carries the complete metadata block, every following line one row array. The streaming end of the artoo workflow (spec -\> apply_spec -\> write_ndjson) for datasets too large for the array-form `.json` file; a thin wrapper over [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) with `format = "ndjson"`. ## Usage ``` r write_ndjson( x, path, on_invalid = c("error", "translit", "fold", "replace", "ignore"), created = NULL, strict = FALSE ) ``` ## Arguments - x: *The dataset to write.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), carrying `artoo_meta`. - path: *Destination `.ndjson` path.* `: required`. A `.ndjson.gz` path writes gzip-compressed bytes. - on_invalid: *Policy for values that are not valid UTF-8.* `: default "error"`. One of `"error"` (abort with `artoo_error_codec`), `"replace"` (substitute `?` and warn with `artoo_warning_encoding`), `"ignore"` (drop the invalid bytes), or `"translit"` / `"fold"` (accepted for pipeline symmetry; behave as `"error"` here, since a byte-level invalidity has no character fold). See [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) for when this fires. - created: *Creation timestamp.* ` | NULL`. `NULL` (default) stamps the current time into `datasetJSONCreationDateTime`; freeze it for byte-stable output. - strict: *Suppress the `_artoo` extension block.* `: default FALSE`. See [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md): the same extension semantics apply to the metadata line. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. ## Details **Bounded memory, both directions.** The writer streams slabs of per-column JSON literals and [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md) parses slab-sized line batches, so a multi-million-row dataset never materializes a whole `rows` array the way the `.json` codec must. A `.ndjson.gz` path gzips the stream transparently. ## See also [`read_ndjson()`](https://vthanik.github.io/artoo/reference/read_ndjson.md) for the inverse; [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) for the array-form file; [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: write a conformed dataset as NDJSON ---- # # apply_spec() attaches the metadata; write_ndjson() streams the metadata # line and one row per line. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".ndjson") write_ndjson(adsl, path) readLines(path, n = 2)[2] #> [1] "[\"CDISCPILOT01\",\"01-701-1015\",\"1015\",\"701\",\"701\",\"Placebo\",\"Placebo\",0,\"Placebo\",0,19725,19906,182,0,0,63,\"<65\",1,\"YEARS\",\"WHITE\",1,\"F\",\"HISPANIC OR LATINO\",\"Y\",\"Y\",\"Y\",\"Y\",\"Y\",\"Y\",null,null,null,25.100000000000001,\"25-<30\",147.30000000000001,54.399999999999999,16,18382,43.899999999999999,\">=12\",19718,\"2014-01-02\",\"2014-07-02\",12,19906,\"COMPLETED\",\"Completed\",23]" # ---- Example 2: gzip the stream via the file extension ---- # # A .ndjson.gz path compresses transparently; read_ndjson() inflates it. gz <- tempfile(fileext = ".ndjson.gz") write_ndjson(adsl, gz) nrow(read_ndjson(gz)) #> [1] 60 ``` # Write a dataset to Apache Parquet Serialize a data frame to an Apache Parquet (`.parquet`) file, storing the data natively while preserving the full `artoo_meta` as a CDISC-shaped sidecar in the file's key-value metadata. The emit end of the artoo workflow (spec -\> apply_spec -\> write_parquet); a thin wrapper over [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) with `format = "parquet"`. Requires the lightweight `nanoparquet` package. ## Usage ``` r write_parquet( x, path, encoding = NULL, on_invalid = c("error", "translit", "fold", "replace", "ignore"), compression = "snappy" ) ``` ## Arguments - x: *The dataset to write.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), carrying `artoo_meta`. - path: *Destination `.parquet` path.* `: required`. - encoding: *Source charset to record.* ` | NULL`. The parquet bytes are always written as UTF-8 (the format's STRING type is UTF-8 by spec); `encoding` only records the data's original charset in the `artoo_meta`, so a later [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) can reproduce the source bytes. `NULL` (default) leaves the recorded encoding untouched. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. - on_invalid: *Policy for values that are not valid UTF-8.* `: default "error"`. One of `"error"` (abort with `artoo_error_codec`), `"replace"` (substitute `?` and warn with `artoo_warning_encoding`), `"ignore"` (drop the invalid bytes), or `"translit"` / `"fold"` (accepted for pipeline symmetry; behave as `"error"` here, since a byte-level invalidity has no character fold). See [`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md) for when this fires; parquet STRING bytes are UTF-8 by spec, exactly like Dataset-JSON. - compression: *Column compression codec.* `: default "snappy"`. One of: - `"snappy"` (default) — fast, the parquet ecosystem default. - `"gzip"` — smaller files, slower. - `"zstd"` — the best size/speed trade-off where supported. - `"uncompressed"` — raw pages. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. ## Details **Metadata where plain Parquet has none.** A bare nanoparquet/arrow file drops labels, formats, and codelists; `write_parquet()` embeds the complete `artoo_meta` as a single Dataset-JSON-shaped string under the `metadata_json` key, so [`read_parquet()`](https://vthanik.github.io/artoo/reference/read_parquet.md) restores every CDISC attribute. The same string is what a `.json` file or an rds carries, so conversion between any two formats stays lossless. A reader without artoo still opens the data and can see the `metadata_json` block. ## See also [`read_parquet()`](https://vthanik.github.io/artoo/reference/read_parquet.md) for the inverse; [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: write a conformed dataset to Parquet ---- # # apply_spec() attaches the metadata; write_parquet() stores the data # natively and the metadata as a CDISC-shaped sidecar. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".parquet") write_parquet(adsl, path) # ---- Example 2: round-trip and confirm the metadata survived ---- # # Reading it back yields an identical artoo_meta. back <- read_parquet(path) identical(get_meta(back)@columns, get_meta(adsl)@columns) #> [1] TRUE ``` # Write a dataset to rds Write a data frame to an R `.rds` file, preserving its `artoo_meta`. A thin wrapper over [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) with `format = "rds"`; the rds carries the metadata both as live R attributes and as the language-agnostic `metadata_json` string, so [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) restores it exactly. ## Usage ``` r write_rds(x, path, encoding = NULL) ``` ## Arguments - x: *The dataset to write.* `: required`. - path: *Destination `.rds` path.* `: required`. - encoding: *Source charset to record.* ` | NULL`. rds is R-native and faithful: strings are saved as-is, never transcoded. `encoding` only records the data's original charset in the `artoo_meta`, so a later [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md) can reproduce the source bytes. `NULL` (default) leaves the recorded encoding untouched. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. ## See also [`read_rds()`](https://vthanik.github.io/artoo/reference/read_rds.md) for the inverse; [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec(cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists) # ---- Example 1: write a conformed dataset to rds ---- # # apply_spec() attaches the metadata; write_rds() carries it into the file. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".rds") write_rds(adsl, path) # ---- Example 2: round-trip and confirm the metadata survived ---- # # Reading it back yields an identical artoo_meta. back <- read_rds(path) identical(get_meta(back)@columns, get_meta(adsl)@columns) #> [1] TRUE ``` # Write a specification to native JSON or a P21 Excel workbook Serialise a `artoo_spec`, dispatching on the file extension: a `.json` path writes artoo's native, lossless JSON; a `.xlsx` path writes a Pinnacle 21 (P21) style Excel workbook. Both are inverses of [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) on their format, which makes the spec converters free compositions: `read_spec("define.xml") |> write_spec("spec.xlsx")` is a Define-XML to P21 bridge in one line. ## Usage ``` r write_spec(spec, path) ``` ## Arguments - spec: *The specification to serialise.* `: required`. Build one with [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md) or [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md). - path: *Destination file.* `: required`. The extension picks the format: `.json` (native, lossless) or `.xlsx` (P21 interchange; needs the `writexl` package). Any other extension aborts with `artoo_error_input`. ## Value *The output `path`, invisibly.* Read it back with [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md). ## Details **Native JSON is the lossless format.** Each slot is written as an array of row objects, with `NA` encoded as JSON `null` and numbers at full precision, so [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) rebuilds an identical `artoo_spec` through [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md). Object keys are emitted in a fixed order, so writing the same spec twice yields byte-identical output. **P21 xlsx is the interchange format.** Sheets are emitted with the headers the P21 reader recognises (Define, Datasets, Variables, ValueLevel, Codelists, Methods, Comments, Documents; empty optional sheets are omitted), foreign keys repeated on every row (no merged cells), and the spec's [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md) as the Datasets sheet's `Standard` column. The study row writes back as the Define sheet's Attribute/Value pairs (`StudyName`, `StudyDescription`, `ProtocolName`). The `Data Type` column is written in the Define-XML / ODM vocabulary the workbook expects: a character variable is `text` (not the Dataset-JSON `string`), and `decimal` / `double` collapse to `float`, `boolean` / `URI` to `text`. Columns the P21 vocabulary does not model are not lost: a foreign column carried on a slot is re-emitted verbatim under its own header, so an xlsx round-trip keeps user columns. **Note:** fields with no P21 column (`itemoid`, `target_data_type`, per-variable `key_sequence`) do not survive an xlsx round-trip; persist to JSON when you need the spec back exactly. The `Data Type` re-encoding is also non-injective: `decimal`, `double`, `boolean`, and `URI` fold to `float` or `text` on a read-back. A Define-XML `partialDate` / `partialDatetime` (and the other partial / incomplete subtypes) is read as the base `date` / `datetime` – CDISC Dataset-JSON v1.1 has no partial dataType – so it is written back as the base type. ## See also **Inverse:** [`read_spec()`](https://vthanik.github.io/artoo/reference/read_spec.md) reads native JSON, a P21 Excel workbook, or Define-XML back into a `artoo_spec`. **Build / inspect:** [`artoo_spec()`](https://vthanik.github.io/artoo/reference/artoo_spec.md), [`spec_datasets()`](https://vthanik.github.io/artoo/reference/spec_datasets.md), [`spec_variables()`](https://vthanik.github.io/artoo/reference/spec_variables.md), [`spec_standard()`](https://vthanik.github.io/artoo/reference/spec_standard.md). ## Examples ``` r # ---- Example 1: persist a spec to JSON, then read it back ---- # # Build a spec from the bundled CDISC-pilot tables, write it to a temp # JSON file, and confirm read_spec() reconstructs it intact. spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) path <- tempfile(fileext = ".json") write_spec(spec, path) identical(read_spec(path), spec) #> [1] TRUE # ---- Example 2: the same spec as a P21 workbook ---- # # The .xlsx path emits P21-shaped sheets; reading the workbook back # recovers the P21-representable surface (here: the dataset names). if (requireNamespace("writexl", quietly = TRUE)) { xlsx <- tempfile(fileext = ".xlsx") write_spec(spec, xlsx) spec_datasets(read_spec(xlsx)) } #> [1] "ADSL" ``` # Write a dataset to SAS XPORT Serialize a data frame to a SAS Transport (`.xpt`) file in v5 (the FDA submission standard) or v8 (extended names and labels), preserving the `artoo_meta` a column can hold. The emit end of the artoo workflow (spec -\> apply_spec -\> write_xpt); a thin wrapper over [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) with `format = "xpt"`. ## Usage ``` r write_xpt( x, path, version = 5, encoding = NULL, on_invalid = c("error", "translit", "fold", "replace", "ignore"), created = NULL ) ``` ## Arguments - x: *The dataset to write.* `: required`. Typically the output of [`apply_spec()`](https://vthanik.github.io/artoo/reference/apply_spec.md), carrying `artoo_meta`. - path: *Destination `.xpt` path.* `: required`. - version: *XPORT transport version.* `: default 5`. `5` (the FDA standard: names \<= 8 characters, labels \<= 40 bytes) or `8` (names \<= 32, long labels). - encoding: *Target charset.* ` | NULL`. `NULL` (default) inherits the source encoding recorded in `artoo_meta`, else UTF-8. IANA and SAS names (`"US-ASCII"`, `"wlatin1"`) both work. **Tip:** any SAS or IANA spelling listed by [`artoo_encodings()`](https://vthanik.github.io/artoo/reference/artoo_encodings.md) is accepted. - on_invalid: *Policy for values not representable in `encoding`.* `: default "error"`. The same policy vocabulary as the UTF-8 writers ([`write_json()`](https://vthanik.github.io/artoo/reference/write_json.md), [`write_ndjson()`](https://vthanik.github.io/artoo/reference/write_ndjson.md), [`write_parquet()`](https://vthanik.github.io/artoo/reference/write_parquet.md)): - `"error"` `(default)` — abort with `artoo_error_codec`, naming the offenders. - `"translit"` — fold smart punctuation (curly quotes, en/em dashes, ellipsis, bullet) to its exact ASCII form per the SAS NLS punctuation table and warn; a character with no fold (a diacritic) still aborts. - `"fold"` — `"translit"` plus the ICU Latin-ASCII accent strip (`Ö` to `O`, `ß` to `ss`, `Æ` to `AE`), and warn. Lossy on names — the original characters are not recoverable; a character neither table maps (the Euro sign) still aborts. - `"replace"` — substitute one `?` per unrepresentable character and warn with `artoo_warning_encoding`. - `"ignore"` — drop the unrepresentable characters silently. **Tip:** for a US-ASCII submission write, `"translit"` fixes the word-processor punctuation that dominates real findings while keeping genuine data corruption loud; reach for `"fold"` only when accent stripping is an accepted, documented step of the migration. - created: *Header timestamp.* ` | NULL`. `NULL` (default) stamps the current time; freeze it for byte-stable output. ## Value *The input `x`*, invisibly, so a write can sit mid-pipeline. ## Details **What XPORT can carry.** An `.xpt` file's NAMESTR stores only variable name, label, length, and SAS format. CDISC metadata beyond that (keySequence, codelist, origin, targetDataType, ...) and the source encoding are not representable in the bytes; they ride the in-session `artoo_meta` and the sidecar in self-describing formats (Dataset-JSON, Parquet, rds). XPORT also cannot distinguish an empty string from `NA` (both store as blanks) and drops trailing spaces. **Character ISO dates (`--DTC`) write as text.** A character column whose `dataType` is `date`/`datetime`/`time` with no numeric `targetDataType` is the CDISC ISO 8601 text form — the SDTM `--DTC` convention — and stores as a character variable, partial dates (`"1951"`, `"1951-12"`) included, byte for byte. The SAS-numeric encoding (with `DATE9.`-style formats) is used for columns that are R `Date`/`POSIXct`/`hms` or whose metadata records `targetDataType = "integer"` (the ADaM numeric-date convention). A character column *under* `targetDataType = "integer"` aborts loudly — a partial date can never become a SAS numeric silently. ## See also [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) for the inverse; [`write_dataset()`](https://vthanik.github.io/artoo/reference/write_dataset.md) for the generic dispatcher. ## Examples ``` r spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) # ---- Example 1: write a conformed dataset as v5 (FDA standard) ---- # # apply_spec() attaches the metadata; write_xpt() carries the label, length, # and SAS format for each variable into the transport file. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") path <- tempfile(fileext = ".xpt") write_xpt(adsl, path) # ---- Example 2: v8 for long names, with a frozen timestamp ---- # # Version 8 keeps names over 8 characters; a fixed `created` makes the bytes # reproducible. Reading it back shows the labels, types, and record count # survived the transport. DM is SDTM, so it conforms against the bundled # sdtm_spec. dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. path8 <- tempfile(fileext = ".xpt") write_xpt(dm, path8, version = 8, created = as.POSIXct("2020-01-01", tz = "UTC")) #> Warning: Widened 1 column past the declared spec length: "STUDYID (7 -> 12)". #> ℹ Values need more bytes than the spec length; data was kept whole. #> ℹ Update the spec length, or shorten the data, so the file matches its #> declared metadata. get_meta(read_xpt(path8))@dataset$records #> [1] 60 ``` # List the members of a SAS XPORT transport file Report every dataset (member) a SAS Transport (`.xpt`) file holds, with its label, variable count, and row count — the survey step before [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) with `member =` picks one. A single-member file (the FDA submission convention) returns one row. ## Usage ``` r xpt_members(path) ``` ## Arguments - path: *Source `.xpt` path.* `: required`. A file that is not a valid XPORT library aborts with `artoo_error_codec`. ## Value *A ``* with one row per member and columns `member` (1-based index), `name`, `label`, `nvars`, and `nobs`. Pass `member` or `name` to [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md). ## Details **v5 has no recorded row count.** A v8 member records its rows; a v5 member's count is derived from the byte span up to the next member (or end of file) minus trailing padding, so an all-character v5 member whose last row is entirely blank reports one row fewer (the documented v5 ambiguity, see [`write_xpt()`](https://vthanik.github.io/artoo/reference/write_xpt.md)). ## See also [`read_xpt()`](https://vthanik.github.io/artoo/reference/read_xpt.md) with `member =` to read one of them. ## Examples ``` r spec <- artoo_spec( cdisc_adam_datasets, cdisc_adam_variables, codelists = cdisc_codelists ) # ---- Example 1: a single-member file reports one row ---- # # The FDA convention is one dataset per transport file. dm <- apply_spec(cdisc_dm, sdtm_spec, "DM", conformance = "off") #> 1 variable the spec declares is absent from the data (not added): #> `BRTHDTC`. p <- tempfile(fileext = ".xpt") write_xpt(dm, p) #> Warning: Widened 1 column past the declared spec length: "STUDYID (7 -> 12)". #> ℹ Values need more bytes than the spec length; data was kept whole. #> ℹ Update the spec length, or shorten the data, so the file matches its #> declared metadata. xpt_members(p) #> member name label nvars nobs #> 1 1 DM Demographics 25 60 # ---- Example 2: survey a multi-member file, then read one member ---- # # Concatenate two single-member files into one library and list it. adsl <- apply_spec(cdisc_adsl, spec, "ADSL", conformance = "off") p2 <- tempfile(fileext = ".xpt") write_xpt(adsl, p2) multi <- tempfile(fileext = ".xpt") writeBin( c( readBin(p, "raw", file.size(p)), readBin(p2, "raw", file.size(p2))[-(1:240)] ), multi ) xpt_members(multi) #> member name label nvars nobs #> 1 1 DM Demographics 25 60 #> 2 2 ADSL Subject-Level Analysis Dataset 48 60 ``` --- name: artoo description: > Read and write clinical-trial datasets losslessly across SAS XPORT, CDISC Dataset-JSON v1.1, Apache Parquet, and RDS from R. Use when writing R code that uses the artoo package. license: MIT compatibility: Requires R >=4.3. --- # artoo Read and write clinical-trial datasets losslessly across SAS XPORT (v5/v8), CDISC Dataset-JSON v1.1, NDJSON, Apache Parquet, and RDS from R. One canonical metadata model (`artoo_meta`) is carried by every codec, so any-to-any conversion is lossless by construction. Pure R, no Java, no SAS, and no compiled runtime beyond a small Parquet engine. ## Installation ```r install.packages("artoo") # development version: # pak::pak("vthanik/artoo") ``` ## Mental model The workflow is one line of three verbs: **read a spec**, **apply it** to a raw data frame, then **read or write** any format. Each verb returns an immutable `artoo_spec` or a conformed `data.frame` that carries the spec's metadata. Nothing is silently transformed at the edges — write is lossless by inheriting the source encoding; ASCII is a gate, not a silent recode. Every condition is a typed `artoo_error_` / `artoo_warning_` / `artoo_message_` so a caller catches by class, not by message. ```r library(artoo) dm <- apply_spec(cdisc_dm, sdtm_spec, dataset = "DM") write_xpt(dm, tempfile(fileext = ".xpt")) ``` Every reader takes `encoding =` (`read_xpt()` auto-detects `windows-1252` when bytes are not valid UTF-8); every writer takes `on_invalid = c("error", "translit", "fold", "replace", "ignore")`, ordered least to most lossy: refuse (default), fold smart punctuation to exact ASCII per the SAS NLS table, also strip accents per ICU Latin-ASCII (`Ö` to `O`, `ß` to `ss`; the BASECHAR analogue), one `?` per unrepresentable character, or drop. `translit`/`fold` still abort on a character with no standards-backed ASCII form (the Euro sign). ## API overview ### Specs Build a spec from metadata tables, or read one from native JSON, a Pinnacle 21 xlsx workbook, or Define-XML; write it back, or amend it in R when the data disagrees. - `artoo_spec`: Build and validate a CDISC spec from metadata tables - `read_spec`: Read a spec from native JSON, a Pinnacle 21 xlsx workbook, or Define-XML - `write_spec`: Write a spec to JSON (lossless) or xlsx (interchange) - `set_type`: Retype a variable in R when the data disagrees with the spec's dataType - `repair_spec`: Flip every `integer_fraction` / `integer_overflow` variable to `float` from a study findings frame, in one call ### Spec accessors Read one slot of a spec. - `spec_standard`, `spec_study`, `spec_datasets`, `spec_variables`, `spec_codelists`, `spec_keys`, `spec_methods`, `spec_comments`, `spec_documents` ### Conform Apply the spec to a raw frame, decode single variables through its codelists, and read or replace the metadata a conformed frame carries. - `apply_spec`: Coerce, order, sort, and stamp a raw frame to its spec (`extra = "drop"` trims to the spec's columns; `on_coercion_loss = "keep"` preserves data an `integer` dataType cannot hold — the divergence is reported, not silently truncated) - `decode_column`: Map a coded variable through a spec codelist - `get_meta`, `set_meta`, `sync_meta`: Read, replace, or refresh the `artoo_meta` a conformed frame carries ### Check Surface conformance findings for one dataset or a whole study, plus the spec's own structural integrity, with the control object that scopes both. - `check_spec`: Conformance findings for one dataset (needs the data) - `check_study`: Conformance findings across a whole study in one pass; prints a dataset-by-check count matrix; feeds `repair_spec()` - `validate_spec`: Structural integrity of the spec itself (no data) - `conformance`: Read the findings `apply_spec()` attached to a frame - `artoo_checks`: Toggle which conformance dimensions run (only `integer_fraction` / `integer_overflow` are fatal coercion checks; `type_mismatch` is informational; `invalid_encoding` flags character bytes that are not valid UTF-8, the signature of a mis-declared source encoding) ### Read and write (lossless any-to-any) Generic dispatch on the file extension, plus a short wrapper per format. - `read_dataset`, `write_dataset`: Generic, dispatch on the file extension - `read_xpt`, `write_xpt`: SAS XPORT v5/v8 - `read_json`, `write_json`: CDISC Dataset-JSON v1.1 (array form) - `read_ndjson`, `write_ndjson`: CDISC Dataset-JSON v1.1 (newline-delimited) - `read_parquet`, `write_parquet`: Apache Parquet (via `nanoparquet`) - `read_rds`, `write_rds`: R serialization ### Inspect The SAS PROC CONTENTS variable pane and the dataset inventory of a file. - `columns`: Variable-pane summary of a conformed frame - `members`, `xpt_members`: Dataset inventory of a file or library ### Reference tables Reference tables for the codecs this session can read and write, and the encoding names R, SAS, and Python share. - `artoo_formats`: Codecs available in this session (extension, direction, engine) - `artoo_encodings`: Charset names across IANA, R, SAS, and Python ### Predicates Class checks for the artoo S7 objects. - `is_artoo_spec`, `is_artoo_meta`, `is_artoo_checks` ### Bundled demo data CDISC pilot specs and datasets rebuilt from public sources. Use these instead of inventing toy data. - `adam_spec`, `sdtm_spec`: pilot ADaM and SDTM specs - `cdisc_adam_datasets`, `cdisc_adam_variables`, `cdisc_sdtm_datasets`, `cdisc_sdtm_variables`, `cdisc_codelists`: the metadata tables the bundled specs are built from - `cdisc_adsl`, `cdisc_adae`: ADaM demo frames - `cdisc_dm`, `cdisc_vs`, `cdisc_ts`, `cdisc_suppdm`: SDTM demo frames ## Conventions (don't fight these) - **Type model is CDISC Dataset-JSON v1.1 verbatim** — `dataType` in {string, integer, decimal, float, double, boolean, date, datetime, time, URI}, `targetDataType` in {integer, decimal}, plus length / displayFormat / keySequence. Dates and times round-trip via dataType + targetDataType + displayFormat, never a class attribute. - **Encoding follows global standards** — IANA names (`US-ASCII`, `windows-1252`, `UTF-8`; SAS `WLATIN1` maps to `windows-1252`), Unicode NFC (UAX #15). Regulatory defaults: xpt = US-ASCII (FDA TCG), json = UTF-8 (CDISC / RFC 8259). - **Metadata rides in the file** — one CDISC-shaped `metadata_json` sidecar is embedded verbatim in every container (parquet KV, rds attr). That is what makes any-to-any lossless: the full itemGroup travels with the data. - **Conformance findings are data, not conditions** — `apply_spec()` attaches a findings tibble read back by `conformance()`; `check_spec()` and `check_study()` return one. Feed the study-level findings straight into `repair_spec()` to fix the spec in one call. - **`extra = "drop"` in `apply_spec()` trims to exactly the spec's columns** — the intended shape for submission; the default `"keep"` passes foreign columns through untouched. - **Errors are classed cli conditions** — `artoo_error_` where kind is one of `input`, `spec`, `type`, `codelist`, `codec`, `validation`, `conformance`. Catch by class, never by message. ## Resources - [Full documentation](https://vthanik.github.io/artoo/) - [llms.txt](https://vthanik.github.io/artoo/llms.txt) — Indexed reference for LLMs - [llms-full.txt](https://vthanik.github.io/artoo/llms-full.txt) — Full documentation for LLMs