ETL tools move data out of the systems that produce it and into the warehouse that reports on it.
This guide ranks them on the things that decide whether the pipeline still runs in a year: how they handle a source that changes its schema, what happens when a load fails halfway through, and how the bill behaves at volume.
Vendors can pay for visibility on this page. It never changes what an entry
says about a product, including the criticism, and we earn nothing when you click through to a
vendor. How that works.
In short
What ETL and data integration software does
ETL and data integration software extracts records from source systems, reshapes them into a common structure, and loads them into a warehouse, database or other application.
In this order: setup effort, what it really costs, how your data comes back out, whether
you can leave, and who each ETL and data integration tool is built for. Why those five, and why there is no
score out of ten, is on the how we work page.
Consumption units (IPUs), quoted; annual commitment; free tier up to 20M rows a month
Yes
Large regulated estates with mainframes and strict lineage requirements
Cloud-native teams with a warehouse and SaaS sources
Country is where the vendor is headquartered or contracts from, which is a
different question from where your data is hosted. Where the two tell different stories, the
entry says so.
Open source connectors you can host yourself and repair yourself
Ranked #1 of 12 in Best ETL Tools in 2026.
Free tierOpen sourceSelf-hostablePublished pricingNorth America
Airbyte is the honest answer to the exit question, because the connectors are source code you already have. When a niche source breaks, you are not waiting on a vendor's backlog.
The catalogue's size is a headline number that hides a distribution: the top fifty connectors are well maintained and the tail is not, so check the specific ones you need for recent commits and open issues. Self-hosted, it needs monitoring, upgrades and someone who understands the scheduler.
What stands out
Open source
Self-hosted
700+ connectors
Where it costs you
Community connector quality varies widely
Self-hosting is real operational work
Right for
Teams with engineers who want to own and modify their pipelines
Wrong for
Business teams with nobody to run the platform
United StatesFree self-hosted open source (Airbyte Core); cloud by volume or credits from $20 a month, published; capacity plans quoted
Czech data platform billed by the minute of processing
Ranked #2 of 12 in Best ETL Tools in 2026.
Free tierPublished pricingEurope
Keboola solves the problem of the team that has a warehouse, a handful of sources and no data engineer. Orchestration and lineage come with the pipelines, the catalogue comes with the enterprise plan, the free plan runs in the EU, and its 60 minutes a month cover a small pilot.
The dependency it creates is broader than a pipeline tool's. If your transformations are written in Keboola and your schedules run there, then leaving means rebuilding the whole layer, not swapping a connector.
What stands out
EU region available
Consumption billing
Orchestration included
Where it costs you
Transformations live inside the platform, which complicates leaving
Minute-based consumption is hard to estimate in advance
Right for
Small European teams wanting extraction and orchestration in one place
Wrong for
Organisations that want the warehouse to be the only platform
CzechiaFree plan with 60 processing minutes a month; top-ups at $0.14 a minute, published; enterprise quoted
Pipelines priced per flow, so the bill stops moving
Ranked #3 of 12 in Best ETL Tools in 2026.
Published pricingNorth America
The per-flow price is the argument. Every other vendor here bills something that grows when your business grows, which turns a data budget into a variable cost nobody controls. Dataddo charges for the pipe, not the water.
It grew up on advertising and SaaS sources where volumes are modest, and it now also replicates databases, where a flat price per flow flatters a busy table. Test that replication at your own volume first, and expect the thin transformation layer to send you to the warehouse anyway.
What stands out
Published price
Predictable unit
Hybrid deployment
Where it costs you
Basic transformation capabilities
Database replication is newer ground than its SaaS connectors; test it at your volumes
Right for
Marketing and finance teams needing predictable monthly integration costs
Wrong for
Teams wanting heavy transformation inside the integration tool
United StatesPer data flow per month from $99, published; 30-day money-back guarantee, no free plan
Danish ELT with SQL transformations and lineage included
Ranked #4 of 12 in Best ETL Tools in 2026.
Published pricingEurope
Weld's proposition is that one product gets a small company from scattered sources to reportable data, with the transformations, scheduling and lineage in the same place and prices you can read online.
For a company of a hundred people that is usually enough, and the European base helps with procurement. Above that scale, the model graph outgrows the tooling and you begin wanting dedicated transformation tooling, at which point you are paying Weld mostly for extraction.
What stands out
EU vendor
SQL transformations
Published price
Where it costs you
Shorter connector list than the larger vendors
Transformation layer strains on complex model graphs
Right for
Analytics teams building a first warehouse without hiring engineers
Wrong for
Very high volume replication or complex modelling
DenmarkMonthly tiers by monthly active rows and connectors from $99, published; 14-day trial
Log-based change capture that also backfills the history
Ranked #5 of 12 in Best ETL Tools in 2026.
Published pricingNorth America
Estuary reads the database write-ahead log, so it captures every change without polling and without hammering the source with queries. Backfill and live tailing are the same pipeline, which removes the usual awkward switchover. Ask honestly whether you need it.
Most reporting is fine on a nightly schedule, and streaming means monitoring lag, handling replay and understanding what happens when the consumer falls behind. Buy it for operational use cases, not because real time sounds better.
What stands out
Change data capture
Streaming
Published price
Where it costs you
Fewer connectors than the established vendors
Streaming adds operational complexity over nightly batch
Right for
Teams needing warehouse data within seconds rather than hours
Wrong for
Reporting needs that a nightly load already satisfies
United States$0.50 per gigabyte plus $100 per connector a month, published; free up to 10 GB a month
Transformation pushed down into the warehouse you already pay for
Ranked #6 of 12 in Best ETL Tools in 2026.
Pricing on requestEurope
Pushdown is the right architecture: the warehouse is already sized and paid for, so running transformation anywhere else duplicates compute. Matillion adds a visual layer over it that analysts genuinely use, which widens who can maintain the pipeline.
The economics need watching. You pay Matillion credits for the duration of the run and the warehouse for the compute, so an unoptimised job is billed by two suppliers, and nobody notices until the quarterly invoice.
What stands out
Pushdown transformation
Warehouse native
Credit pricing
Where it costs you
Credits are consumed by job runtime, so slow jobs cost twice
Visual pipelines are awkward to review in version control
Right for
Warehouse-centred teams wanting visual transformation analysts can read
Wrong for
Engineering teams who prefer transformation as plain code
United KingdomCredits consumed by task hours; annual credit packages, rate quoted
Managed connectors that keep working without anyone watching them
Ranked #7 of 12 in Best ETL Tools in 2026.
Free tierPublished pricingNorth America
Fivetran sells the absence of a problem: connectors that survive source API changes, schema drift that lands as new columns rather than as a failed load, and no maintenance work for your team. That is worth real money and it is why it stays on shortlists.
The pricing unit is the risk. Monthly active rows count any row touched, so a table with constant updates costs the same as one with constant inserts, and forecasting means running a real load first.
What stands out
Managed connectors
Schema drift handling
Row pricing
Where it costs you
Monthly active row pricing is volatile and hard to predict
A base charge per connection adds up across many small sources
Right for
Teams who want connectors that simply keep working, and will pay
Wrong for
Budget-sensitive teams with frequently updated tables
United StatesPer monthly active row, published calculator; free plan up to 500,000 rows a month
French-founded integration engine with master data and quality attached
Ranked #8 of 12 in Best ETL Tools in 2026.
Self-hostablePricing on requestNorth America
Semarchy comes from the classic integration world, where the hard part was never the connector but the reconciliation: which customer record is the real one, and what to do with the three variants.
Bundling that with the movement of data makes sense for manufacturers and public bodies with systems older than the cloud. It is the wrong shape for a startup with a warehouse and twelve SaaS sources, and the buying process assumes a project rather than a card payment.
What stands out
French origin
On-premise option
Master data
Where it costs you
Quoted pricing with an implementation partner expected
Tooling feels older than the SaaS competition
Right for
Enterprises integrating on-premise systems with master data rules
Wrong for
Small teams loading SaaS sources into a cloud warehouse
United StatesQuoted per organisation; SaaS, on-premise or inside Snowflake; trial available
The old Talend stack inside a larger analytics company
Ranked #9 of 12 in Best ETL Tools in 2026.
Open sourcePricing on requestNorth America
The data quality tooling is the reason to consider it: profiling, standardisation and rules that produce evidence an auditor accepts, sitting in the same product as the pipelines. For a regulated enterprise that is a real consolidation.
The history is the problem. Organisations that built on the open source Studio have had to plan migrations they did not choose, and the current product is priced and packaged for large accounts rather than for the teams that adopted Talend originally.
What stands out
Data quality
Wide connectivity
Enterprise contract
Where it costs you
Migration from legacy Talend Studio projects is difficult
Roadmap now follows Qlik's analytics priorities
Right for
Enterprises needing data quality and profiling alongside their pipelines
Wrong for
Teams that adopted open source Talend and want continuity
United StatesQuoted capacity subscription, metered on volume, job runs and duration
The pipeline runner already inside your Azure agreement
Ranked #10 of 12 in Best ETL Tools in 2026.
Self-hostablePublished pricingNorth America
Data Factory wins on procurement and on reach: it is already in the agreement, and the self-hosted integration runtime pulls from databases behind the firewall without opening anything to the internet. That combination is hard to beat for a hybrid estate.
Day to day it is unpleasant. Pipelines are configured through a slow interface, failures are investigated in run histories, and because each activity run is billed, teams cram logic into fewer activities and make the pipelines harder to follow.
What stands out
Azure native
Per-run pricing
Hybrid runtime
Where it costs you
Authoring and debugging experience is poor
Per-activity pricing distorts pipeline design
Right for
Azure estates moving data between on-premise systems and the cloud
Wrong for
Teams needing a broad catalogue of SaaS application connectors
United StatesPer activity run and per integration unit hour, published
Serverless Spark jobs, drawn visually or written in code
Ranked #11 of 12 in Best ETL Tools in 2026.
Published pricingNorth America
Glue is infrastructure rather than a product, and priced accordingly: per second of processing capacity, with no licence sitting on top. For heavy transformation over files in object storage, nothing here is cheaper.
It assumes a team that already writes Spark, tunes partitioning and reads driver logs. It also does not solve the boring half of the job, which is extracting from a hundred SaaS APIs, so most estates end up pairing Glue with something from higher on this page.
What stands out
Serverless Spark
AWS native
Visual or code
Where it costs you
SaaS coverage limited to a few enterprise applications
Requires Spark skills to use properly
Right for
AWS data teams comfortable writing and tuning Spark jobs
Wrong for
Loading marketing and SaaS sources without engineering
United StatesPer data processing unit hour, published; billed per second, 1-minute minimum
The enterprise incumbent, priced in units nobody forecasts
Ranked #12 of 12 in Best ETL Tools in 2026.
Free tierSelf-hostablePublished pricingNorth America
Informatica remains the answer when the sources are old, the auditors are serious and the estate is too large for a tool a team maintains itself. Connectivity to mainframe and legacy applications is real and rare.
The costs are equally real: consumption units that bundle unrelated activities, an annual commitment negotiated in advance, and engineers who charge more because the skill is scarce. Before signing, price the same three pipelines on two other tools and put the difference in the paper.
What stands out
Enterprise governance
Mainframe connectivity
Consumption pricing
Where it costs you
Consumption unit pricing is difficult to forecast
Needs specialist skills that command a premium
Right for
Large regulated estates with mainframes and strict lineage requirements
Wrong for
Cloud-native teams with a warehouse and SaaS sources
United StatesConsumption units (IPUs), quoted; annual commitment; free tier up to 20M rows a month
ETL and data integration software extracts records from source systems, reshapes them into a common structure, and loads them into a warehouse, database or other application. The differences that matter are rarely in the feature list, so this is
the order we would work through them.
01
Decide whether you need a published price
9 of the 12 tools here publish what they cost; the other 3 quote per organisation. The ones you can compare without a sales call: Airbyte, Keboola, Dataddo, Weld, Estuary, Fivetran, Azure Data Factory, AWS Glue, Informatica IDMC (Salesforce).
02
Decide how much the jurisdiction matters
These 12 vendors are established in 4 countries across 2 regions (North America 9, Europe 3). That decides whose disclosure law applies to what the vendor holds, wherever the servers are.
03
Consider whether you want the source
2 of these are open source: Airbyte, Qlik Talend Cloud. Hosting one yourself trades a subscription for maintenance.
Transform before the warehouse, or inside it
The old pattern transformed data in a dedicated engine and loaded the result: Informatica IDMC and Qlik Talend Cloud grew up this way, and it still suits estates where the destination is not one big warehouse. The current pattern loads raw data first and transforms it with the warehouse's own compute, which is what Fivetran and Airbyte assume and what Matillion turns into a visual product by pushing the work down into Snowflake or BigQuery.
ELT is usually cheaper, because you already pay for that compute, and easier to debug, because the raw data is still there to compare against. It also means your warehouse bill absorbs the transformation cost, so watch both invoices together.
Decide where transformation runs before comparing connector lists.
If you choose ELT, model the extra warehouse compute in the same business case.
Keep raw loaded data for long enough to reprocess a bad transformation.
Batch or streaming, and whether you actually need seconds
A nightly batch is simple, cheap and sufficient for most reporting. Change data capture, which Estuary does as its core function and Fivetran and Airbyte offer for databases, reads the transaction log and delivers changes continuously. It puts far less load on the source than repeated polling, which matters for a production database that also serves customers.
The cost is operational: you now monitor lag, handle replays after an outage, and reason about what a consumer sees mid-transaction. Choose streaming for operational uses, fraud checks, stock levels, service dashboards. Choose batch for finance and management reporting, and do not let a vendor sell you seconds you will never look at.
Ask which business decision changes if data is an hour old instead of a minute.
Check what change capture does to your source database's log retention settings.
Test a replay after a deliberate outage before committing to streaming.
Schema drift and the load that fails halfway
Connector counts are marketing; the real question is what happens on a Tuesday when a source adds a column or renames one. Good tools land the new column in the warehouse and carry on. Weaker ones fail the whole load, or quietly drop the field, which is worse. Fivetran handles this better than anything else here and charges for it.
The second question is idempotency: if a load dies at seventy percent, does rerunning it duplicate rows or resume cleanly. Ask about the state the pipeline keeps between runs, and about how deleted source records are represented, because a soft delete that never reaches the warehouse produces reports that quietly overstate everything.
During the trial, add and rename a column in a source and watch what arrives.
Kill a running load and rerun it; count the rows afterwards.
Ask how deletions in the source are reflected in the destination.
Rows, connectors or compute: three bills that behave differently
Fivetran counts monthly active rows, so a table updated constantly costs the same as one growing constantly, and a chatty source can double your bill without adding information. Dataddo charges per flow, which is flat and predictable but assumes modest volumes.
Keboola bills processing minutes, AWS Glue and Azure Data Factory bill compute, and Matillion bills credits by runtime, so all four reward efficient jobs and punish careless ones. Informatica IDMC uses consumption units that mix several activities into one currency. Airbyte self-hosted charges nothing and bills you in staff time instead. Run a real month of your own data through the two finalists before signing anything annual.
Load one real month of data during the trial and read the resulting bill.
Ask what an update to an existing row costs compared with an insert.
Check the price of a full historical resync, which you will need at least once.
What goes wrong most often when buying ETL and data integration software
Choosing on connector count. You need eleven connectors and only their quality matters; the other four hundred are for someone else's stack.
Sizing the contract on today's volumes. Row-based pricing follows business growth, and the renewal arrives after the growth, not before it.
Assuming a failed load can simply be rerun. Ask about resumption and duplicates before you find out during a month-end close.
Putting transformation logic somewhere it cannot be reviewed. If pipelines are not in version control, nobody can say what changed when the numbers moved.
07
Frequently asked questions
8 answers
What is the best ETL and data integration in 2026?
Airbyte leads our ranking of 12. The only entry here you can run free on your own servers, with connector source you can read and fix when it breaks, which is the independence argument in one sentence.
The catalogue passes 700 connectors because anyone may contribute, and that is the weakness too: community connectors range from production-grade to abandoned. Self-hosting is genuine operational work, the platform is licensed under ELv2 rather than a classic open source licence, and the cloud pricing model has been revised more than once.
Which ETL and data integration tools publish their pricing?
9 of the 12, with the pricing model each one publishes:
Airbyte: Free self-hosted open source (Airbyte Core); cloud by volume or credits from $20 a month, published; capacity plans quoted.
Keboola: Free plan with 60 processing minutes a month; top-ups at $0.14 a minute, published; enterprise quoted.
Dataddo: Per data flow per month from $99, published; 30-day money-back guarantee, no free plan.
Weld: Monthly tiers by monthly active rows and connectors from $99, published; 14-day trial.
Estuary: $0.50 per gigabyte plus $100 per connector a month, published; free up to 10 GB a month.
Fivetran: Per monthly active row, published calculator; free plan up to 500,000 rows a month.
Azure Data Factory: Per activity run and per integration unit hour, published.
AWS Glue: Per data processing unit hour, published; billed per second, 1-minute minimum.
Informatica IDMC (Salesforce): Consumption units (IPUs), quoted; annual commitment; free tier up to 20M rows a month.
The other 3 quote per organisation.
Is there a free ETL and data integration tool?
Airbyte, Keboola, Fivetran, Informatica IDMC (Salesforce) offer a free tier or a free self-hosted edition.
Where are these ETL and data integration vendors established?
In 4 countries across 2 regions: North America 9, Europe 3.
Airbyte: United States.
Keboola: Czechia.
Dataddo: United States.
Weld: Denmark.
Estuary: United States.
Matillion: United Kingdom.
Fivetran: United States.
Semarchy: United States.
Qlik Talend Cloud: United States.
Azure Data Factory: United States.
AWS Glue: United States.
Informatica IDMC (Salesforce): United States.
Which ETL and data integration tools are open source?
Airbyte, Qlik Talend Cloud.
Which ETL and data integration tools can you host yourself?
Airbyte, Semarchy, Azure Data Factory, Informatica IDMC (Salesforce). The other 8 are hosted by the vendor only.
What should you use instead of Airbyte?
Keboola and Dataddo are the next two on this page. Keboola is for small European teams wanting extraction and orchestration in one place; Dataddo is for Marketing and finance teams needing predictable monthly integration costs.
Who should not buy Airbyte?
Business teams with nobody to run the platform. Community connector quality varies widely.
If your ETL and data integration product belongs among these 12, tell us what it does and who it is for. Inclusion is an editorial call; what a listing is and is not is set out under software advice.