ISSUE 2026-09-27SERIES BREC 37TYPE BRIEFgarnetgrid.com
The actual question
Insights
The first data hire usually gets made to relieve pain that analysts have been absorbing for a year, which means the job spec gets written from the pain rather than from the work. That produces a list of fourteen technologies and no decision about what the person is for. Here is what the role actually is, how to tell whether you need it yet, and an interview that separates people who can make a number trustworthy from people who can name tools.
Most companies phrase this as "we need a data engineer". The more useful version is narrower: which specific datasets does the business now make decisions on, and who is accountable for those being correct tomorrow morning when nobody is watching? If you can name three and nobody owns them, you have a data engineering problem and you should hire. If you cannot name three, you have a different problem, and your first data engineer will spend their first quarter discovering that instead of fixing it.
What the first one is actually for
Not a platform. Resist the platform. The first data engineer's job is to take a small number of datasets and make them repeatable, inspectable and trustworthy: the same inputs produce the same outputs, the lineage is written down, and something alarms before a human notices the number is wrong.
You are ready when analysts spend more time reconciling than analysing. When two dashboards disagree and the explanation lives in someone's head. When a report breaks silently and the discovery mechanism is a person in a meeting on Thursday. When the load runs on a laptop, or on a cron job on a machine nobody can name.
You are not ready if you have three sources and one weekly report that someone runs by hand. That is an ownership problem, not an engineering one. A scheduled job with a named owner and an alert on failure costs a day of work. Hiring against it costs a year, and the person you hire will be bored by March.
Three jobs advertised as one
Analytics engineering is modelling and semantics: the definitions, the tests on them, the dbt-shaped layer that turns raw tables into something a finance person can use. It sits closest to the business.
Data engineering proper is ingestion, orchestration, schema and contract management, backfills, and being on call for the pipes.
Data platform work is the substrate: storage layout, compute, access control, cost, upgrades, whatever cluster you are running.
Most first-hire job specs are the union of all three, plus some machine learning for luck. Pick one as the centre of gravity. My view: if your data mostly lives in SaaS tools and one transactional database, hire analytics-leaning and buy the ingestion. If you are running on hardware you own, or your volumes mean you are building ingestion rather than buying it, you need the middle profile with genuine infrastructure depth, and you should expect that to be harder to find. The overlap between people who can read a Postgres query plan and people who can debug a systemd unit is not where most CVs sit.
Hire for the failure modes, not the tool list
A job spec listing fourteen technologies tells you the hiring manager has not decided what the work is. The work is mostly the following, and it recurs everywhere regardless of stack.
A candidate who raises several of these unprompted has done the job. A candidate who lists tools may only have watched it being done.
The interview that actually discriminates
Do not run an algorithms round. It tells you nothing about whether someone can make a number trustworthy.
Instead, hand them a wrong number. Two tables, a job log, and a dashboard showing revenue that fell last week when you know it did not. Ask them how they would find out where it went. Strong candidates ask when the number changed before theorising about why, ask for the log before the schema, and go straight to the boundary between two systems, because that is where data goes missing.
Then ask the real question: how would you know this table was wrong before the CFO told you? You are listening for reconciliation against an independent source, row-count and distribution checks, freshness thresholds, and above all a check that is capable of failing. Ask about a backfill that went badly and what they changed afterwards. Ask what they would delete from a stack they inherited, because a willingness to remove things is rarer and more useful than enthusiasm for adding them.
The red flags are consistent. Every problem is answered with a new tool. Nothing they built ever woke them up, which usually means they never operated it. They cannot describe a rollback. And the tests they are proudest of cannot fail: checking that a table is not empty, on a table that has never once been empty, is a control that validates nothing.
What has to be true before they start
Access on day one. Credentials, a machine, the warehouse, a production replica, the repository. Someone who spends three weeks in a procurement queue draws the correct conclusion about how serious you are.
One named problem with a named beneficiary. "Make the weekly revenue figure reconcile to the payment processor, and keep it reconciled" is a job. "Modernise our data stack" is a wish with a budget attached.
One decision made for them, and one decision kept away from them. Made for them: where compute and data live. Kept away: what the business terms mean. If sales and finance disagree about what an active customer is, no engineer resolves that. They will encode one side of the argument and then be blamed by the other.
Managed platforms or your own hardware
Decide this before you write the job spec, because it changes the hire.
On cost, the honest answer is that I cannot tell you which is cheaper, and neither can anyone who has not looked at your volumes, your retention and your query patterns. What is knowable is the shape of the bill. Managed warehouses meter consumption: credits, DBUs, bytes scanned, tokens for anything model-shaped, seats for the tooling on top. Cost therefore tracks behaviour, and one badly written dashboard refreshing every five minutes becomes a standing charge nobody authorised. Owning the hardware converts that into capital up front plus the salary cost of operating it: capacity planning, upgrades, backups that are actually restored rather than merely scheduled, and somebody reachable at night.
Both models share an invariant worth saying out loud. Schema evolution, data quality and cost discipline are your work either way. Managed services remove the machines, not the engineering.
There is also a reason to own the hardware that has nothing to do with cost. If the data cannot leave your perimeter, because of regulation, a customer contract, or simply because you do not want records of real people sitting in a third party's account, the decision is already made and the arithmetic is irrelevant. Be clear which reason you are acting on, because the two lead to different hires. A managed stack lets you hire closer to the business. Your own hardware means your first data engineer is also part site-reliability engineer, and if you pretend otherwise the backups will not be tested.
Do it in this order. Write down the three datasets the business genuinely decides on, and the name of the person accountable for each. If the list is empty, or every name is the same overworked analyst, fix that before you open a role. Decide where compute and data live, and be honest about whether the reason is cost or perimeter. Then write the spec from the failure modes rather than a tool list, interview with a wrong number and a job log, and have the credentials issued before their first morning. Success for a first data engineer fits in one sentence: the numbers the business relies on are right, and when they stop being right you hear it from a check rather than from a customer.