Skip to content

Systems and records

API rate limits and full-history exports: how long a large pull takes

By SourceX Editorial · Updated

Short answer

A large API export takes roughly the total number of calls divided by the calls allowed per minute. Total calls are records divided by page size, plus extra calls for child records such as comments, plus attachments and retries. Estimate each object separately, use bulk or incremental endpoints where offered, and plan around the slowest object.

Key takeaways

  • Export time is total calls divided by allowed calls per minute, plus waits for retries and limit resets.
  • Per-parent calls for comments, reviews or activities usually dominate the estimate, not the parent list.
  • Bulk and incremental endpoints can change an estimate more than any amount of parallel code.
  • Measure real throughput on a sample before committing to a schedule.
  • Production integrations may share the same limit, so heavy pulls belong outside business hours.

The basic formula for export time#

The basic formula for API export time is the number of calls divided by the calls allowed per minute. The number of calls is the record count divided by the page size, plus any extra calls for child records, attachments and retries.

The inputs come from two places. Record counts come from the system itself, usually through reports or admin screens. Page sizes and rate limits come from the vendor's developer documentation and your plan, and they can differ by endpoint, app type and plan tier.

  • List calls = records divided by records returned per page.
  • Child calls = one or more calls per parent record, unless the API returns children in bulk.
  • Minutes of calling = total calls divided by calls allowed per minute.
  • Elapsed time = minutes of calling plus waits for limit resets, retries and any windows when the pull must pause.

Which inputs change the estimate most?#

The inputs that change an export estimate most are record counts per object, page sizes, the binding rate limit and attachment sizes. Get the counts first: ask the system owner for records per object and per year before any engineer writes code.

When a vendor sets both a short-window limit and a daily cap, work out the time under each and use the longer result. A daily cap means the pull cannot finish faster than total calls divided by the daily allowance, however quickly each minute runs. That longer figure is the one to give finance and leadership when they ask when the export will be done.

Which inputs change the estimate most?
InputWhere to find itCommon trap
Record count per objectAdmin reports, database counts or the vendor's usage pagesCounting parents only and forgetting comments, activities or line items
Records per pageEndpoint documentationAssuming the largest page size applies to every endpoint
Calls allowed per minute or per dayRate limit documentation and your planIgnoring daily caps that bind long before per-minute limits
ConcurrencyRate limit and app documentationRunning parallel workers against one shared limit
Shared quotaThe account's other integrationsProduction integrations using the same allowance during business hours
Attachment sizeStorage reportsTreating file downloads as free when they are often the slowest part

Why child records multiply the call count#

Child records multiply the call count because many APIs return a parent list cheaply but need a separate request for each parent's details. A help desk may list tickets in large pages but return comments one ticket at a time. A code platform may list pull requests in pages but need further calls for each one's reviews, review comments and timeline.

This pattern turns a pull that looks short on paper into a long one. When a system holds many more child records than parents, per-parent calls dominate the estimate and the page size of the parent list barely matters.

Look for endpoints that return children in bulk, such as incremental exports of all comments by date, or options that include related records in the parent response. Finding one is usually worth more than tuning the code.

Bulk API versus REST: which to use#

A bulk or asynchronous export API is usually the better choice for a full-history pull when the vendor offers one, and paginated REST is the fallback for objects bulk endpoints do not cover. Many large platforms offer both, with separate limits.

Most real pulls mix methods: bulk jobs for large flat objects, REST for child details and attachments, and an incremental pass at the end to catch records that changed while the job ran.

Bulk API versus REST: which to use
MethodHow it worksGood forWatch for
Paginated RESTRequest pages until none remainMedium objects, child details, objects without bulk supportPer-call limits, deep pagination caps, per-parent calls
Bulk or asynchronous exportSubmit a job, wait, download result filesLarge flat objects such as accounts, orders or history tablesJob queue time, file size limits, fields bulk jobs exclude
Incremental or cursor exportPull everything changed since a timestamp or cursorVery large objects and resumable pullsDeleted records, cursor expiry
Vendor full-account exportAdmin-triggered archive of the accountA baseline copy before cancellationFormats, missing attachments, export frequency caps
Database or warehouse syncReplicate tables continuouslyOngoing history once connectedBackfilling history from before the sync began

What do published limits look like in practice?#

Published export limits come in several shapes: calls per hour, exports per minute, items per export and exports per week. Each shape constrains the plan differently, which is why one formula needs different inputs per system.

The examples below are taken from vendor documentation and change over time. Use them to recognize the type of limit you face, then confirm the current figure for your plan before you build the schedule.

What do published limits look like in practice?
SystemDocumented limitEffect on a full-history pull
GitHub REST API5,000 requests per hour with a personal access token; 15,000 per hour for a GitHub App installation on Enterprise Cloud, with secondary limits applied separatelyPer-pull-request calls for reviews and comments set the pace, so the credential type matters
GitLab project file exportDefault of 6 project exports per minute and 1 export download per minute per userMany repositories mean queueing downloads; admins on self-managed instances can change the defaults
Jira Cloud CSV exportUp to 10,000 work items per asynchronous export, with larger sets split into JQL batches, and one CSV export processed at a timeBatch by date range and run batches in sequence
Salesforce Data Export ServiceA manual export once every 7 days on Enterprise, Performance and Unlimited, or every 29 days on Professional and DeveloperA missed or incomplete run can cost a week or a month, so plan API pulls alongside it

Illustrative: estimating a help desk pull#

Illustrative: a fictional vertical software company wants every ticket and comment from its help desk before a platform change. The figures below are invented to show the method; they are not any vendor's limits.

The archive holds 400,000 tickets. The list endpoint returns 100 tickets per page, so listing takes 4,000 calls. Comments come back one ticket at a time, which adds 400,000 calls. At an allowance of 200 calls per minute, listing takes about 20 minutes and comments take about 2,000 minutes, more than 33 hours of uninterrupted calling before retries or attachments.

The engineer then finds an incremental endpoint that returns comments for all tickets by date. With about four comments per ticket and 1,000 comments per page, the comment pull drops to about 1,600 calls. The figures are made up, but the shape is typical: per-parent calls set the timeline, and the right endpoint matters more than faster code.

What else stretches a large pull#

Several factors beyond the published rate limit stretch a large pull, and most appear only once the job is running. Build for them from the start.

Checkpoint progress so a failed job resumes from the last completed window rather than starting over. Store raw API responses as well as transformed files, so a later mapping error does not force a second pull.

  • Limit responses: back off and retry when the API signals a limit, and log every retry.
  • Shared limits: production integrations may draw on the same allowance, so schedule heavy phases off-hours.
  • Deep pagination: some APIs cap how far offset paging goes, so split the pull into date windows or use cursors.
  • Token expiry: long jobs need credentials that refresh without manual steps.
  • Attachments: file downloads can carry separate limits and much larger payloads.
  • Changes during the pull: records edited mid-pull need a final incremental pass.

How SourceX scopes a large export#

SourceX scopes a large export during Supply, the first step of the SourceX five-step transaction, using metadata the company already has: systems, object counts, years of history and known export routes. No data moves during that assessment.

If a package goes ahead, the export runs in the company's environment and the data stays in its own storage or ships on encrypted drives. The SourceX Evidence Packet records provenance, including the export method and date, so a buyer can see how the dataset was produced.

Frequently asked questions

Can we ask the vendor to raise our API limit for a one-time export?

It is worth asking. Some vendors grant temporary increases, run a bulk export on request or point you to an endpoint built for migrations. Ask through your account team with a clear description of the objects, counts and date range, and get any increase confirmed in writing.

Is it faster to export from a database backup?

For systems you host yourself, usually yes, because no API limit sits between you and the tables. SaaS products rarely give database access, though some vendors offer warehouse sync or data sharing options. Check whether those include history tables and attachments before relying on them.

Do vendor rate limits change over time?

Yes. Vendors revise limits, retire endpoints and apply different rules to different app types, sometimes with little notice. Re-read the documentation right before the pull, and build the job to read limit headers and adapt rather than assume a fixed rate.

Should we run a sample before the full pull?

Always. A sample of one object over a narrow date range shows real throughput, response sizes and error rates. Multiply from measured numbers, not from documentation, and you will catch per-parent calls and pagination caps before they stall the full job.

Will a large pull disrupt our production integrations?

It can if they share an allowance. Check whether the limit applies per app, per user or per account, use a dedicated app or credential where possible, and run the heaviest phases outside business hours so customer-facing integrations keep working.

Sources

  • GitHub's REST API allows 5,000 requests per hour with a personal access token and 15,000 per hour for a GitHub App installation on GitHub Enterprise Cloud; secondary rate limits also apply. Source
  • GitLab's default rate limits for file exports are 6 project exports per minute per user and 1 project export download per minute per user. Source
  • Atlassian supports exporting up to 10,000 work items with the asynchronous Export CSV feature and recommends splitting larger exports into JQL batches. Source
  • Atlassian's cloud changes notes say only one Jira CSV export can be processed at a time. Source
  • Salesforce's Data Export Service allows a manual export every 7 days in Enterprise, Performance and Unlimited Editions, and every 29 days in Professional and Developer Editions. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify