Why data warehouse automation on top of Azure Data Factory?
Azure Data Factory is the execution layer. It copies data, it orchestrates and it runs on a schedule, and it does that well. What ADF does not do is decide which tables need loading today, how you recognise a change, where the history lives, or what happens when a source system renames a column. You build that yourself, or you have it generated. Yres does the second: it writes the ADF pipelines, the tables and the load logic from metadata, inside the customer's own Data Factory.
Last updated:
What does Azure Data Factory do on its own?
ADF gives you the connectors, the copy activity, the orchestration and the triggers. You point at a source and a target, and it moves the data. It records per run whether a pipeline succeeded.
That gives you an engine, not a data warehouse. ADF does not know what a customer record is, or that branch three has booked differently since the merger, or that last month's address change has to be kept. You put that knowledge in, and that is the work.
What do you build on top of it?
With one source it is manageable. With thirty sources and a few hundred tables it becomes a second product you maintain alongside your data warehouse.
Microsoft documents a metadata-driven pattern for ADF itself: a control table holding the objects and load parameters, and parameterised pipelines that take their work from it. That is exactly the right direction. It is the start of the list, not the end of it.
| Component | Does ADF provide it? | What it means in practice |
|---|---|---|
| Connecting a source | Yes, connector available | Authentication, pagination and error handling are configured per source |
| Fetching source metadata | No | You build the retrieval and tracking of columns and types yourself |
| Staging | No | Your own layer, your own naming, your own cleanup |
| Incremental loading | Partly | A watermark or CDC has to be devised and tracked per source |
| Change detection | No | Hashes or comparisons written by hand |
| History (SCD2) | No | The merge logic is the hardest part you build yourself |
| Deleted records | No | A record that disappears from the source disappears silently |
| Schema changes | No | A renamed column looks like a removed plus a new column to a warehouse |
| Promotion to test and production | Partly | ADF has git and publish; which changes belong together it does not |
| Checks on the content | No | A green pipeline says nothing about whether the numbers are right |
| Log retention | No | Log tables keep growing until somebody intervenes |
What does Yres generate instead?
Everything in that table, from one metadata model. You configure a source and the tables you want, and Yres writes the ADF pipelines and linked services into your own Data Factory, with git integration to your own Azure DevOps. The tables and load procedures land in your own Azure SQL.
Seven load types cover the ways a source can behave, from a full copy to a change feed. Six of them build history on arrival, because most SaaS APIs return only the current state and the history therefore cannot come from there.
Connecting a new table after that is configuration, not development. That is the whole point: the difference between source one and source one hundred should be as small as possible.
- ADF pipelines and linked services, in your own Data Factory
- Staging and history tables, with SCD2 without writing the merge
- Incremental loading with the watermark that suits the source
- A change process across DTAP: what belongs together moves together
- Monitoring per load and per step, plus checks on the content rather than only on the run
What Yres deliberately does not do
The data model. Kimball, Data Vault or a variant of your own: that choice depends on your organisation and your people, not on your tooling. Yres delivers the layer any method can sit on and leaves the modelling to the team.
That is an architectural choice rather than an omission, but it is one you have to want. If you are looking for a product that generates a dimensional model for you, other vendors serve you better and we say so.
When you are better off keeping ADF by hand
There are situations where an automation layer costs more than it returns.
The tipping point is not a particular number of sources but how much a new source costs once the first one works. If that stays an afternoon, nothing is wrong. If source twenty-one costs as much as source one, you are paying for your standardisation in hours instead of in licence.
- A handful of sources that rarely change
- A strong engineering team that wants to work code-first and maintain the standardisation itself
- An architecture built around Databricks or Snowflake rather than Azure SQL
- Streaming or near-real-time as a core requirement; Yres loads on a schedule
What stays yours
Everything sits in your own Azure subscription: the pipelines in your own Data Factory, the tables and history in your own Azure SQL, your secrets in your own Key Vault, the git history in your own Azure DevOps.
Pipelines that Yres generates are also managed by Yres: manual edits to them are overwritten on the next generation. Whatever you put in the custom folder always stays, through upgrades as well. You outsource the standard work and keep the exceptions.
Frequently asked questions
Does Yres replace Azure Data Factory?
No. Yres uses ADF as its execution layer and generates the pipelines that run in it. You keep working on Azure Data Factory, you just stop building the pipelines by hand.
Can I keep my existing ADF pipelines?
Yes. Whatever sits in your Data Factory's custom folder is left alone by Yres, through regeneration and after upgrades. Your own pipelines run alongside the generated ones.
Does Yres generate a data model?
No, deliberately not. Yres delivers the ingestion and history layer; the dimensional or Data Vault model you build yourself with your own views and procedures. If you want a tool that produces that model for you, other products fit better.
What happens to the ADF pipelines if I stop using Yres?
They stay. They run in your own Data Factory, in your own subscription, with git history in your own Azure DevOps. We impose no restrictions when you leave.
How is this different from Microsoft's own metadata-driven copy?
That pattern solves the copying: a control table drives parameterised pipelines. Yres does that too, and additionally the history, the change detection, the promotion across DTAP, the checks on the content and the maintenance. The difference is everything after the copy.
