knitr::opts_chunk$set(
collapse = TRUE,
comment = "#>",
warning = FALSE,
message = FALSE
)
library(ctOpenData)
library(dplyr)
library(ggplot2)
Welcome to the ctOpenData package, an R package designed
to provide convenient access to the Connecticut Open Data Portal.
The package provides a streamlined interface for discovering and
downloading datasets from Connecticut Open Data. It helps bridge the gap
between raw Socrata API endpoints and tidy data analysis in
R.
The package provides three primary functions:
ct_list_datasets() for browsing available datasetsct_pull_dataset() for downloading datasets using a
catalog key or Socrata UIDct_any_dataset() for downloading data directly from a
Socrata JSON endpointThe first step in a typical workflow is to use
ct_list_datasets() to retrieve the live Connecticut Open
Data catalog.
catalog <- ct_list_datasets()
catalog
The returned catalog includes information about the datasets available through the portal. Two especially important columns are:
key, a human-readable dataset identifier generated from
the dataset nameuid, the official Socrata dataset
identifierYou can search the catalog for datasets containing a keyword.
catalog |>
filter(grepl("spill", name, ignore.case = TRUE)) |>
select(key, uid, name)
Replace KEYWORD with a useful search term related to the
example dataset selected for the package.
The primary way to download data is with
ct_pull_dataset().
A dataset can be requested using either its human-readable catalog
key or its official Socrata UID.
example_data_uid <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 5
)
example_data_uid
example_data_key <- ct_pull_dataset(
dataset = "spill_incidents_from_july_1_2022_to_recent_for_download",
limit = 5
)
example_data_key
Both calls should return data from the same dataset.
Dataset keys are easier to read, while Socrata UIDs are
more stable.
For reproducible research and long-term workflows, using the official
Socrata UID is generally recommended.
The filters argument can be used for simple exact-match
filtering.
filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 25,
filters = list(
incident_type = "Petroleum Incident"
)
)
filtered_data
You can confirm that the filter worked by inspecting the unique values in the selected field.
filtered_data |>
distinct(incident_type)
Multiple values can also be supplied.
filtered_multiple <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 50,
filters = list(
incident_type = c("Biomedical Incident", "Dielectric Fluid Incident")
)
)
filtered_multiple
Multiple fields can be combined within the same filter list.
filtered_combination <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 50,
filters = list(
incident_type = "Petroleum Incident",
township = "New Haven"
)
)
filtered_combination
If the example dataset contains a date or datetime field, records can
be filtered using from, to, and
date_field.
date_filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
from = "2023-01-01",
to = "2024-01-01",
date_field = "reported_date",
limit = 100
)
date_filtered_data
The from date is inclusive, while the to
date is exclusive.
A single day can also be requested using the date
argument.
single_day_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
date = "2023-01-01",
date_field = "reported_date",
limit = 100
)
single_day_data
Socrata EndpointThe preferred workflow is to use ct_list_datasets()
together with ct_pull_dataset().
However, when a dataset is not available in the package catalog,
ct_any_dataset() can download data directly from a
Socrata JSON endpoint.
Connecticut Open Data endpoints typically follow this structure:
https://data.ct.gov/resource/<dataset_uid>.json
For example:
https://data.ct.gov/resource/ffju-s5c5.json
The endpoint can then be supplied directly to
ct_any_dataset().
endpoint_data <- ct_any_dataset(
json_link = "https://data.ct.gov/resource/ffju-s5c5.json",
limit = 5
)
endpoint_data
Use ct_pull_dataset() when the dataset is available
through ct_list_datasets().
Use ct_any_dataset() when you already have a valid
Socrata JSON endpoint or when the dataset is not included
in the package catalog.
Once the data have been downloaded, they can be analyzed using standard R tools.
The following example counts the number of records in a categorical field.
category_summary <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 500
) |>
filter(!is.na(incident_source)) |>
count(incident_source, sort = TRUE)
category_summary
The results can then be visualized.
category_summary |>
slice_head(n = 10) |>
ggplot(
aes(
x = n,
y = reorder(incident_source, n)
)
) +
geom_col() +
theme_minimal() +
labs(
title = "Most Frequent Categories",
x = "Number of Records",
y = "Category"
)
This example demonstrates the complete workflow from discovering a dataset to downloading, filtering, summarizing, and visualizing it.
The ctOpenData package provides a consistent interface
for working with data from the Connecticut Open Data Portal.
In this vignette, you learned how to:
ct_list_datasets()ct_pull_dataset()Socrata JSON endpoint using
ct_any_dataset()These functions allow users to focus on analysis rather than manually constructing API requests.
If you use this package for research or educational purposes, cite it using the package citation returned by:
citation("ctOpenData")