Skip to main content
Connectors

Apify Dataset connector

Set up the Apify Dataset connector in Kaivo: authentication, configuration, the 4 BigQuery tables it syncs, and answers to common questions.

Written By Lauri Raivio

Last updated 16 days ago

Kaivo is a fully managed data platform that syncs your Apify Dataset data into a Google BigQuery warehouse and keeps it up to date automatically. There is no pipeline to build and no infrastructure to run, so you can spend your time analysing your data from Apify Dataset instead of moving it.

What is the Apify Dataset connector

Sync your Apify datasets and scraped items into BigQuery with Kaivo to analyse crawl results and track output volume across your scraping runs.

CategoryFiles & Databases, Tech
AuthenticationAPI key
SetupSelf-service

Getting started with the Apify Dataset connector

  1. Sign up for Kaivo and create a workspace.
  2. Connect your Apify Dataset account.
  3. Choose which tables to sync.
  4. Wait for the initial sync to finish.
  5. Query your data in BigQuery or your favourite AI or BI tool.

Authenticating Apify Dataset

Authenticate with your API Token.

FieldDescription
API Token

Your personal Apify API token. Find it in the Apify Console under Settings > Integrations.

Configuring the Apify Dataset connector

When you set up the connector, you provide:

FieldDescription
Dataset ID

ID of the Apify dataset to load. Find it in the Apify Console under Storage > Datasets.

Tables and columns synced from Apify Dataset

Kaivo syncs 4 tables from Apify Dataset into a dedicated dataset in your BigQuery warehouse. Click any table to see its columns and types.

ColumnTypeDescription
_kaivo_idSTRINGPrimary key that uniquely identifies the row. Auto-generated by Kaivo.
idSTRINGUnique identifier of the dataset
nameSTRINGName of the dataset
user_idSTRINGUser ID of the owner of the dataset
created_atSTRINGTimestamp when the dataset was created
stats__read_countFLOAT64Number of times the dataset was read
stats__storage_bytesFLOAT64Total storage size of the dataset in bytes
stats__write_countFLOAT64Number of times the dataset was written to
modified_atSTRINGTimestamp when the dataset was last modified
accessed_atSTRINGTimestamp when the dataset was last accessed
item_countFLOAT64Total number of items in the dataset
clean_item_countFLOAT64Number of clean items in the dataset
act_idSTRINGIdentifier of the actor associated with the dataset
act_run_idSTRINGIdentifier of the actor run associated with the dataset
titleSTRINGTitle of the dataset
fieldsJSONList of fields available in the dataset
_kaivo_extracted_atTIMESTAMPTimestamp that shows when the row was extracted. Auto-generated by Kaivo.
ColumnTypeDescription
_kaivo_idSTRINGPrimary key that uniquely identifies the row. Auto-generated by Kaivo.
idSTRINGUnique identifier of the dataset collection
nameSTRINGName or title of the dataset collection
user_idSTRINGUser ID of the owner of the dataset collection
created_atSTRINGDate and time when the dataset collection was created
modified_atSTRINGDate and time when the dataset collection was last modified
accessed_atSTRINGDate and time when the dataset collection was last accessed
item_countFLOAT64Total number of items in the dataset collection
usernameSTRINGUsername of the owner of the dataset collection
stats__read_countFLOAT64Number of read operations performed on the dataset collection
stats__storage_bytesFLOAT64Total storage size in bytes occupied by the dataset collection
stats__write_countFLOAT64Number of write operations performed on the dataset collection
schemaSTRINGData schema or structure of the dataset collection
clean_item_countFLOAT64Number of clean items in the dataset collection
act_idSTRINGIdentifier of the actor associated with the dataset collection
act_run_idSTRINGIdentifier of the actor run associated with the dataset collection
titleSTRINGDisplay title of the dataset collection
fieldsJSONFields present in the dataset collection
_kaivo_extracted_atTIMESTAMPTimestamp that shows when the row was extracted. Auto-generated by Kaivo.
ColumnTypeDescription
_kaivo_idSTRINGPrimary key that uniquely identifies the row. Auto-generated by Kaivo.
_kaivo_extracted_atTIMESTAMPTimestamp that shows when the row was extracted. Auto-generated by Kaivo.
ColumnTypeDescription
_kaivo_idSTRINGPrimary key that uniquely identifies the row. Auto-generated by Kaivo.
crawl__depthFLOAT64Depth level of the crawled page
crawl__http_status_codeFLOAT64HTTP status code of the response
crawl__loaded_timeSTRINGTime when the page was loaded
crawl__loaded_urlSTRINGURL of the loaded page
crawl__referrer_urlSTRINGURL of the page that referred to the current page
markdownSTRINGMarkdown content of the webpage
metadata__canonical_urlSTRINGCanonical URL of the webpage
metadata__descriptionSTRINGDescription of the webpage
metadata__language_codeSTRINGLanguage code of the webpage content
metadata__titleSTRINGTitle of the webpage
textSTRINGText content of the webpage
urlSTRINGURL of the webpage
screenshot_urlSTRINGURL of the screenshot of the webpage
_kaivo_extracted_atTIMESTAMPTimestamp that shows when the row was extracted. Auto-generated by Kaivo.

How the Apify Dataset sync works

After the first load, Kaivo keeps your BigQuery warehouse up to date for you. Where Apify Dataset supports it, each sync pulls only new and changed records so it stays fast; otherwise it refreshes the whole table. Every record keeps its original ID, so you won't get duplicate rows.

Frequently asked questions

How long does the initial sync take for Apify Dataset?

It depends on how much history is in your Apify Dataset account. Most initial syncs finish within minutes, while large accounts can take a few hours. After that, syncs only fetch new and changed records, so they're much faster.

Can I sync only some tables or columns?

Yes. You pick which tables to sync when you set up the connection and can change the selection later. Tables you don't select are never copied to your warehouse.

What happens when Apify Dataset's schema changes?

New fields are never added automatically. You choose which fields to sync, so data you haven't selected (sensitive personal data, for example) never lands in your warehouse. When a new field appears, it becomes available for you to add. What happens to removed or renamed fields depends on a table's sync mode: full-refresh tables always match what's currently in Apify Dataset, so dropped fields disappear, while incremental tables keep their existing columns and history, so an old field stays and newly added fields fill in over time.

How do I handle GDPR or data deletion requests?

Your data lives in your own Kaivo-managed BigQuery warehouse, so the most direct option is to delete or anonymise specific records right in BigQuery. If you delete data in Apify Dataset instead, full-refresh tables drop it on the next sync, while incremental tables keep it, so you would remove the row in BigQuery or ask us to run a full refresh. To remove everything, delete the Apify Dataset connector in Kaivo and all of its synced data is deleted with it.

Common use cases for Apify Dataset data

Crawl output

Use dataset stats and itemCount to track how many items each run produces over time.

Scraped content

Bring item_collection_website_content_crawler with url and text into BigQuery to search and analyse crawled pages.

Dataset inventory

Report on dataset_collection by createdAt and user to see what you have collected.

Freshness

Compare modifiedAt and accessedAt to find stale datasets worth refreshing or cleaning up.

Use Apify Dataset data in your AI and BI tools

Once Apify Dataset data lands in your Kaivo-managed BigQuery warehouse, you can explore it with AI tools or any BI tool that connects to BigQuery. Here's how the most common destinations work with Apify Dataset data.

Claude

Use Kaivo's MCP server to give Claude secure, workspace-scoped access to your data. Setup guide →

Power BI

Microsoft's BI tool with a native BigQuery connector. Supports direct query and scheduled refresh. Setup guide →

Data Studio

Free Google BI tool with native BigQuery support. One-click connection to your Kaivo warehouse; great for SMB teams on Google Workspace. Setup guide →

Tableau

The premium analytics standard, with native BigQuery integration. Setup guide →

Google Sheets

Use Connected Sheets to query BigQuery directly from a spreadsheet, with no SQL. Setup guide →

Excel

Connect via Power Query's BigQuery connector. Setup guide →

Metabase

Open-source BI tool with strong BigQuery support. Setup guide →

See our pricing page for Apify Dataset connector pricing and plan details.

Was this helpful?

Still need help? Share an idea