Open data

The open certification dataset

Exam facts for 92 IT certifications across 13 vendors, free and machine-readable, and every fact carries the official page it came from. It is licensed CC BY 4.0, so you can reuse it in a tool, a spreadsheet or an article with a link back as the only condition.

Download the JSON

What the file holds

The file has no questions in it and never will; the corpus is a product, not a dataset. What it publishes is the surrounding fact table that is oddly hard to find anywhere complete, 92 of the entries priced from a vendor page:

  • Exam code and vendor

    The exam code, the vendor, and the level, linked to the certification page.

  • Exam format

    Question count, duration, passing score, and the item formats the real exam uses, including whether it has performance-based items and where that is stated.

  • Cost

    Voucher price, retake cost, a prep-budget range, and how the vendor states renewal, with the source page for each.

  • Domain weightings

    The published domain split, normalised to sum to about 100. A note flags any certification where the vendor publishes no weighting, so an assumed split is never read as an official one.

  • Lifecycle

    Whether the exam is active, retiring or retired, with the vendor-announced end date and the page it was read from. This is the fact vendors publish least consistently.

  • Provenance

    Every fact carries the official page it was taken from, and the file records the day it was generated, so nothing goes stale silently.

License and reuse

Licensed CC BY 4.0. Use it commercially or not, change it, build on it. The one condition is attribution. To cite it:

Certification data from FirstTry (https://firsttry.app), CC BY 4.0.

The file records the day it was generated, and each row keeps the source URL behind its numbers, so anything you build on it can show its work the same way the site does.

How a machine reads it

The dataset is described with schema.org/Dataset on this page, and FirstTry also publishes an llms.txt describing what is worth reading and citing. Both point back to the JSON above.