Project 2: Zero-Emissions Vehicles in California

In this project, you will analyze recent data collected by the State of California on new zero-emissions vehicles. Fundamentally, your goal is to characterize the trends in zero-emissions vehicles over the last few years using summary statistics.

By completing this project, you will:

Part 1: The CEC Dashboard

The California Energy Commission is California’s primary energy policy and planning agency, established by the Legislature in 1974 and headquartered in Sacramento. One of its projects is to track data on the adoption of zero-emission vehicles (ZEVs) across the state. They present a current version of that data in a dashboard:

Read every component of that page to understand the source of the data, then interact with the dashboard to explore the data.

What are three different questions that can be answered by interacting this dashboard? Select three questions that you are truly curious about.

Part 2: The Data Behind the Dashboard

Next, follow the the links to download the data behind the dashboard, specifically a file called New_ZEV_Sales_Last_updated_07-17-2026_ada.xlsx. Save it to your computer in your project repo in a subdirectory called data/. Open the file in Excel if you have it, otherwise you can import it to Google Spreadsheets and view it there.

To read the “County” sheet into R as a data frame, you can use the read_excel() function in the readxl package.

# install.packages("readxl") # run this once to install the package
library(readxl)
zev <- read_excel("data/New_ZEV_Sales_Last_updated_07-17-2026_ada.xlsx",
  sheet = "County")

You can then either print the data frame to the screen to see it or, more practically, use View(zev) to open the Positron data viewer.

After reading all three sheets of the Excel file as well as the documentation on the original dashboard site, state as precisely as possible: does this data record new ZEV sales? Or new purchases? Or new registrations? What types of transactions are included and what are excluded from this tally?

For the Country sheet in particular, what does every row refer to (i.e. what is the unit of observation)?

Part 3: Checking the Dashboard

Find where on the Dashboard each of the following statistics are reported, then recreate them from the zev data using a dplyr pipeline. If there are any you cannot recreate, state why not.

  1. Cumulative light-duty ZEV sales through 2026
  2. Year-to-date 2026 total sales
  3. Year-to-date 2026 total sales for the county that you’re from (if you’re not from California, use Alameda County, where we are now).
  4. The 2026 total sales separated by BEV, PHEV, and FCEV
  5. The year-to-date (YTD) Zev Share
  6. A table of counts of 2026 Tesla sales, split by model
  7. A table of counts of 2026 sales, split first by make, then by model1

Part 4: Beyond the Dashboard

There are many more interesting phenomena that we can learn about using this data set. Select three of the following five questions to answer. For each one, state the strongest reason someone could argue that your analysis is misleading.

A. Your Questions

Return to the questions that you wrote in Part 1 and select one to answer with the data (or if you cannot, explain what additional data you would need to answer it).

B. Tax Credit Effect

The previous federal administration put in place financial incentives to encourage people to purchase electric vehicles. Those incentives ended on September 30, 2025. Is there a Q3 2025 surge followed by a Q4 drop, and does it differ between BEVs and PHEVs?

C. Market Concentration

Compute a Herfindahl index across makes for each year. Is the ZEV market becoming more or less concentrated?2

D. Brand Geography

Which make is disproportionately popular in each county compared with its statewide share? For example, a certain car make might make up only 15% of sales statewide but might be 45% of sales in Alameda. That make would be disproportionately population in Alameda county.

Think carefully about how to construct this statistic. If needed, you are welcome to exclude data if that would make the statistic more meaningful.

E. The Elon Effect

In January and February of 2025, Elon Musk, the CEO of Tesla, took an active role in the newly created Department of Government Efficiency (DOGE). The actions of that agency were politically divisive, leading some left-leaning voters to shun the Tesla brand.

Is there any evidence of this Elon effect in the data? That is: did Tesla sales drop in left-leaning counties relative to right-leaning counties?

Here are two lists of clearly-leaning counties, based on the 2024 presidential election results:

dem_counties <- c("Alameda", "San Francisco", "Santa Clara", "Marin",
                  "San Mateo", "Los Angeles", "Santa Cruz", "Humboldt")
rep_counties <- c("Kern", "Shasta", "Modoc", "Lassen", "Tehama", "Kings",
                  "Placer", "El Dorado")

Some things to think about as you put together your analysis:

  • Which year(s) should you compare 2025 to? Note that Tesla’s share was already changing from year to year before 2025.
  • Should your statistic be Tesla’s raw sales count or Tesla’s proportion of all ZEV sales in a county? Think about what each version does and doesn’t control for.

Footnotes

  1. The dashboard carefully arranges the rows of this table. If you’d like an extra challenge, replicate this arrangement using R.↩︎

  2. Hint: this can be most cleanly done by writing a function that takes a vector of counts and returns the index value. That function can then be applied as a summary statistic inside summarize().↩︎