DataEngineered

Curated, provenance-tracked datasets
built from open public sources

DataEngineered designs and maintains structured, relational datasets across ten consumer and technical verticals — every record traceable to a documented public source, sold as downloadable snapshots, subscriptions, and APIs.

10 dataset products
500K+ curated records
100% provenance-tracked
Monthly refresh pipelines

The datasets

Each product ships a free sample so you can validate schema and quality before licensing.

🧴

INCIdb

Skincare & Cosmetics

Cosmetic formulations decomposed to their INCI ingredient lists, linked to compound-level chemistry.

19,645 formulations44,816 INCI compounds

RoasterDB

Specialty Coffee

Specialty-coffee products with origin, process, and tasting notes mapped to the SCA Flavor Wheel.

8,000+ products280+ roasters20+ countries
💊

SuppDB

Supplements & Nootropics

Real supplement labels from the NIH DSLD, normalized to mg dosages with proprietary-blend flags and PubChem chemistry.

17,000+ products2,000+ brands
🌿

FloraDB

Houseplants & Botany

Individually verified houseplants with quantitative light/water/temperature care metrics, ASPCA pet toxicity, and GBIF-verified taxonomy.

270 curated species891 toxicity records20,000+ species index
🔧

MechanicDB

Automotive OBD-II

Diagnostic trouble codes — SAE-universal plus OEM codes for 32 makes — mapped to ranked repair procedures, DIY difficulty ratings, and aftermarket parts cost estimates.

15,886 DTC codes32 makes (OEM)56,561 ranked fixes75,055 parts mappings
🥃

WhiskyDB

Fine Spirits

Whiskies and fine spirits from government label registries and open registers, with a 19-year monthly auction-price index.

1,290+ spirits3,200+ distilleries20,000+ price benchmarks
🛎️

RecallDB

Product Recalls & Safety

Official U.S. federal product recalls normalized across CPSC, FDA, FSIS, NHTSA and USCG — every record linked back to its source pull with a raw-payload fingerprint.

127,783 recalls292,790 products5 federal sources
🍷

WineDB

Fine Wine & Vintages

Fine wine producers, cuvées and vintages with exact varietal blend percentages, organoleptic tasting descriptors, and secondary-market valuation indices.

3,675 vintages106 wineries57 appellations
🔌

ApplianceDB

Home Appliance Repair

Home-appliance error codes mapped to ranked repair procedures and replacement parts, keyed by composite (brand, appliance type, code) identity across 13 major brands.

438 error codes288 ranked repairs13 brands
🧬

CSA

Clinical-Stage Biotech

Clinical-stage drug assets linked asset → trial → listed sponsor → ticker → FDA status, with a forward catalyst calendar. Every row carries a source URL.

2,268 Phase-3 trials1,841 drug assets2,103 forward catalysts

How we build

The same engineering discipline behind every product.

Open sources only

Government registries, institutional APIs, and open databases — collected politely, within each source's terms, from public endpoints.

Provenance ledgers

Every record carries a source reference: name, endpoint, license, and access timestamp. Any fact can be traced back.

Honest coverage

No invented values. Field-level coverage numbers are published up front; attributes a source doesn't publish stay empty or carry a documented default.

Maintained pipelines

Scheduled refreshes with deduplication keep snapshots current — datasets are living products, not one-off scrapes.

Frequently Asked Questions

Technical facts, provenance verification, and licensing procedures.

How does DataEngineered verify the provenance of its datasets?

Every dataset record includes a transparent provenance ledger containing the exact public source name, API endpoint, license classification, and collection timestamp. We never synthesize or invent missing values.

How frequently are DataEngineered dataset snapshots refreshed?

All ten database products are maintained through automated monthly ingestion pipelines that deduplicate records and append new entries while preserving historical schema compatibility.

Can enterprise teams license customized subsets or tailored data schemas?

Yes. Beyond self-serve Stripe checkout for standard snapshots, enterprise data engineering teams can request custom curation runs on specific target lists or specialized export formats including Parquet, DuckDB, and JSON.

Are free sample data partitions available before purchasing a commercial license?

Absolutely. Every dataset card on the DataEngineered portal links directly to a free sample partition hosted on Hugging Face or Kaggle, allowing your team to test and validate schema compatibility instantly.

Licensing & custom work

Full datasets are sold under commercial licenses — most are self-serve with secure Stripe checkout and instant download, as one-time snapshots or refreshed subscriptions. Custom curation on your target list is available on request.

Direct Email