A painful part of working with your data has always been that your live data and
By Andy Jassy · October 1, 2026 · Curated by George's Blog
A painful part of working with your data has always been that your live data and historical data are stuck in separate systems: the order a customer just placed lives in your database, while their last five years of orders sit in a data lake in S3.
And answering a real question usually needs both at once (is this a normal purchase for them, or should we flag it?), and to do that you had to move the data together first, copying history out of the data lake into your database (or the other way around), because the database couldn’t read it where it lived.
That meant guessing ahead of time which data you’d want, keeping a second copy of it all, building pipelines to move it, and constantly syncing so the two didn’t drift apart. A lot of plumbing, and slow going, all before you could answer one question. And even then, the answers were only as fresh as your last sync.
That now changes with Aurora PostgreSQL, which can call your live data and historical data in S3 together, in a single query. No copying, no pipelines to keep in sync.
And it’s fast, because we’ve built in DuckDB, a popular open source engine that’s really good at reading and analyzing data right where it’s stored. DuckDB reads the open formats like Parquet and Iceberg already sitting in your data lake, so there’s nothing to convert or move.
As folks build AI agents into their apps, the data their agent needs will depend on the task in front of it. Being able to query that specific data live, instead of copying it over just in case, is gonna be a big help for builders. https://lnkd.in/gi76Ux3y