E‑commerce Variant Extraction

Capture every size, colour, and style — not just the parent product

Most scrapers stop at the product title. Our variant‑aware pipelines interact with dropdowns, swatches, and JavaScript selectors to capture every possible combination — per‑variant SKU, price, stock, and image. The result: a complete variant matrix delivered directly into your data warehouse, ready for your pricing engine or PIM.

Variant pipelines trusted by
500M+Variants extracted annually
99.9%Variant capture rate
2–3 wksTypical pipeline delivery
variant-matrix.json
// Complete variant matrix per product
{
  "parent_sku": "SNEAKER-99",
  "variants": [
    { "size": "8", "color": "Black", "sku": "SN99‑8-BLK", "price": 129.99, "stock": 12 },
    { "size": "9", "color": "Black", "sku": "SN99‑9-BLK", "price": 129.99, "stock": 3 },
    { "size": "8", "color": "White", "sku": "SN99‑8-WHT", "price": 129.99, "stock": 0 }
  ]
}
🎨 All combinations captured
24 variants from 4 sizes × 6 colours
📦 Per‑variant stock data
Out‑of‑stock flagged in real time

Variant extraction pipelines running for

SizeWise Retail ColourLab SKUMaster VariantIQ InventoryPulse StyleGrid SizeWise Retail ColourLab SKUMaster VariantIQ
Overview

Every combination. Every stock level. No gaps.

Product variant extraction is the systematic capture of every possible configuration of a product — size, colour, material, style, or any other option — from an e‑commerce product page. While a basic scraper might grab only the default variant, our pipelines interact with every selector, dropdown, and swatch to build the complete variant matrix.

We extract per‑variant SKUs, prices (list and sale), stock availability, and images, mapping everything to your internal product schema. The result is a ready‑to‑use dataset that powers repricing, inventory tracking, and personalised shopping experiences — without the 80% data gap that manual scraping leaves behind.

  • Captures every size, colour, style, material, and configuration
  • Per‑variant SKU, price, sale price, stock, and image
  • Handles dynamic selectors: dropdowns, swatches, and JavaScript
  • Builds the full variant matrix — no missing combinations
  • Delivered in your schema (JSON, CSV, Parquet) to any destination
Business challenges

Why variant data is the hardest part of product scraping

Parent products are easy. The real value — and complexity — lies in the variants hidden behind clicks and swatches.

01

Selectors are invisible to HTTP requests

Variant options like colour swatches or size buttons often inject content via JavaScript. A plain HTTP scraper sees none of this — capturing only the default, leaving 80% of the SKU catalog untouched.

02

Combinatorial explosion

A T‑shirt in 4 sizes and 6 colours yields 24 variants. Manually scraping each combination is impossible. Our pipelines iterate programmatically, capturing every combination — even when some are out of stock and hidden from view.

03

Stock and price differ per variant

The same product can have different prices for different sizes, or stock levels that vary by colour. Without per‑variant capture, you miss critical signals like which specific options are selling out — or being discounted.

Our solution

A variant extraction engine that leaves no combination behind

Every feature is designed to handle the real interactivity of modern e‑commerce variant selectors.

Complete variant matrix iteration

We programmatically click through every option combination (size × colour × style) and capture the resulting SKU, price, and stock. No manual mapping, no missed variants.

Dynamic selector interaction

Headless browsers interact with dropdowns, colour swatches, image tiles, and custom JavaScript pickers — waiting for variant‑specific data to load before extracting.

Per‑variant data enrichment

Beyond SKU, price, and stock, we capture variant‑specific images, weight, dimensions, and any other attributes that differ between options — all mapped to your schema.

Real‑time variant monitoring

After the initial extraction, we set up ongoing monitoring that detects stock level changes, price updates, and new variant additions — pushing alerts and refreshed data on your schedule.

variant-matrix-iteration.json
// Automatic combination loop
{
  "product_url": ".../sneaker-99",
  "options": { "size": [8,9,10], "colour": ["Black","White"] },
  "variants_extracted": 6,
  "status": "complete — 100% coverage"
}
Process

From variant‑heavy pages to a clean matrix feed

1

Variant selector mapping

We reverse‑engineer the site's variant UI — identifying all option types, their possible values, and the underlying JavaScript or API calls.

2

Matrix build & interaction logic

Our engineers build the automation: clicking through every combination, waiting for page updates, and capturing the resulting data for each variant.

3

Sample variant dataset

A subset of products with their full variant matrices is delivered for your validation — you review completeness and schema conformance before full rollout.

4

Full production deployment

All products and all variants go live on your schedule. Continuous monitoring ensures new variants and stock changes are captured automatically.

What you receive

Deliverables at each stage of variant extraction

1
Week 1

Variant selector map & interaction strategy

A document cataloguing every variant selector, option set, and interaction pattern — signed off before development begins.

2
Week 2

Sample variant matrix (500 products)

Variant data for 500 products, capturing all combinations and per‑variant attributes, delivered in your target schema for validation.

3
Week 3

Full‑scale variant feed + runbook

Complete variant data for the entire catalog, with a runbook covering the interaction logic, monitoring, and update cadence.

Ongoing

Variant change detection & weekly health report

Weekly summary of new variants discovered, stock/price changes captured, and any UI modifications handled — proactively shared with your team.

500M+
Variants extracted annually
99.9%
Variant coverage rate
2–3 wks
Average pipeline delivery
from scoping to production
120+
E‑commerce platforms supported
Who needs variant extraction

Teams that operate on per‑SKU, not per‑product, data

👕

Apparel & Footwear

Capture every size‑colour combination with per‑variant stock — critical for inventory planning and availability alerts.

💻

Electronics

Extract configuration variants (RAM, storage, colour) with per‑option pricing and stock — essential for like‑for‑like comparisons.

🛋️

Home & Furniture

Collect fabric, finish, and size options that change price and availability — feed your product configurator with real competitor data.

🧴

Beauty & Cosmetics

Capture shade variations, pack sizes, and gift‑set options — each with its own SKU, price, and stock level.

Why choose managed variant extraction

Variant extraction vs. standard product scraping

Capability Managed Variant Extraction Generic Product Scraping Manual Variant Collection
Full variant matrix (all combinations)
Dynamic selector interaction (JavaScript)
Per‑variant stock & price capture±
Incremental change detection
Ongoing maintenanceIncludedYour teamN/A
Time to complete variant dataset2–3 weeksWeeks of scriptingMonths
What variant users say

“We finally know which specific size is selling out — not just the parent product”

★★★★★

"We spent months trying to scrape variants from a site that loads size charts on click. ScraperScoop delivered the full matrix for 45,000 products in under three weeks — including stock levels for every colour and size. Our inventory team now lives in this data."

SW
Head of Inventory AnalyticsSizeWise Retail
★★★★★

"We needed to compare prices across retailers at the variant level — because different colours often have different discounts. Their pipeline captures that nuance perfectly, and the data drops into Snowflake with the same schema as our internal catalog."

CL
Director of PricingColourLab
★★★★★

"The variant monitoring is what we didn't know we needed. Now we get a Slack alert every time a competitor adds a new size or colour — we can react to assortment changes in hours, not weeks."

SM
Product Intelligence ManagerSKUMaster
Integrations

Variant data lands exactly where your teams need it

Schema‑matched, with per‑variant granularity — feed your pricing engine, inventory system, or PIM directly.

🗄️
Amazon S3
❄️
Snowflake
🔷
BigQuery
🐘
PostgreSQL
🔗
Webhooks (JSON)
📄
CSV / JSON / Parquet

Get a fixed‑price quote for your variant extraction project

Tell us the retailer, the variant types, and your update cadence — we'll return a scope, price, and timeline within two business days.

Frequently asked

Common questions about product variant extraction

Product variant extraction is the process of capturing every possible variation of a product — such as size, colour, material, or configuration — from an e‑commerce product page. We extract per‑variant data including SKU, price, sale price, stock level, image URL, and variant‑specific attributes, and deliver it as a structured, machine‑readable dataset.
Our pipelines use headless browsers to interact with all types of variant selectors — dropdowns, colour swatches, size buttons, and JavaScript‑driven pickers. We wait for the page to load the variant‑specific content before capturing the data, ensuring no variant is missed because it loads on click.
Yes. We programmatically iterate through every possible combination of variant options (e.g., all size‑colour pairs) to build a complete variant matrix. Even when a site shows only one combination at a time, our pipeline captures the full set — down to the last stock‑keeping unit.
For every variant we extract the unique SKU, price (list and sale), stock availability, image URL (if variant‑specific), and any additional attributes such as material finish, weight, or dimensions that differ per variant. The output matches your internal product feed schema.
After the initial full extraction, we set up scheduled or near‑real‑time refreshes that detect changes in price, stock, and new variant additions. Incremental updates push only the differences to your warehouse, reducing load and ensuring your catalog is always current.

Your competitor's variant matrix — delivered as a clean, queryable dataset

Tell us which retailer and which variant types you need, and we'll send a sample variant feed within two business days. No commitment — just proof that complete variant coverage is possible.

Most variant pipelines deliver the full matrix within 2–3 weeks of kickoff.

99.9% variant coverage Per‑variant price & stock Ongoing variant monitoring
From the blog

Insights on retail pricing strategy

🚀 Start Your Data Project

Get a Free Data Sample

See exactly what our data looks like before you commit. No credit card required, no spam — unsubscribe anytime.

  • Custom sample matching your target cuisines and cities
  • Full JSON/CSV export of live menus
  • Dedicated food data expert walkthrough
  • POC turnaround within 24 hours

Start Extracting Data Today

Tell us your requirements and get a custom quote within 2 Working Hours.