Author: MTF Institute Research Team
Independent review: MTF Institute Research QA
Evidence date and geography: 5 October 2026, United States
Technical report number: MTF-CF-RR-2026-10-05-59
Version DOI: 10.5281/zenodo.23171923
Full report and 100-source appendix: Public PDF
Abstract
What does a data engineer do when the intended consumer is a business analyst, decision maker, or operational team? We examined a structured purposive sample of 100 current U.S. employer and employer-authorized job postings, checked on 5 October 2026. The postings describe work that connects source systems to usable analytical data: ingestion, transformation, modeling, quality controls, documentation, monitoring, and interfaces with the people who define and use business measures. They also show meaningful variation. Some roles concentrate on platform and pipeline engineering; others own semantic models, dashboards, or data products close to business decisions. This report interprets those explicit descriptions through linked employer examples. It does not treat a named product as a universal requirement, infer a work rhythm from a job title, or estimate the prevalence of a capability in the U.S. labor market.
Question and method
The research question was: Which responsibilities, work products, methods, qualifications, collaboration patterns, and decision boundaries do current U.S. postings explicitly assign to data engineers building data products for business analytics? The unit of analysis was a distinct public opening, not a company, technology, or search result.
We selected roles with a stated U.S. workplace or explicit U.S. remote eligibility, an active application route, and substantial production engineering work that supported analytical use. Eligible work included source ingestion, transformation, warehouse or semantic modeling, orchestration, quality, or monitoring. An Analytics Engineer title was included only when production data engineering was a principal part of the role. Analyst-only, data-science-only, dashboard-only, and BI-dominant jobs were excluded. The sample contains 89 Data Engineer or Data Engineering titles and 11 Analytics Engineer or Data Analytics Engineer variants. Those numbers describe the selected corpus, not the U.S. occupation's title mix.
The 100 openings come from 80 named employers. Amazon and Walmart contribute five requisitions each after an employer-concentration screen. Fifty-four titles contain a senior, Sr., staff, lead, principal, III, or IV marker. This is a title-string measure: employer levels are not standardized, and the senior tilt limits conclusions about entry-level hiring. Each retained role was checked at its employer or authorized applicant-tracking page. We normalized tracking variants of URLs and compared same-employer, same-title candidates by team, location, salary band, responsibilities, and qualifications so that materially repeated advertisements did not inflate the sample.
For each posting, the coding record distinguishes assigned duties from professional outputs, hard methods from observable collaboration, named tools from tool requirements, and required criteria from preferences. It also records education, experience, title level, expressly stated cadence, interfaces, and decision or escalation language. A short phrase from the individual posting anchors every stated code. If the accessible posting is silent, the field is coded unstated. Silence is not evidence that an activity never occurs. The linked cases below illustrate individual postings. The reviewed matrix counts that follow describe only the selected corpus.
Quantitative summary within the selected postings
The figures below count source-anchored mentions among the 100 selected postings. Each row uses a reviewed category definition built from source-anchored exact codes; n counts postings with at least one included code. A posting can enter more than one row, so the counts are not additive and do not estimate prevalence across the U.S. labor market. Duties and outputs were disclosed in all 100 selected postings, but an unmentioned category is not automatically known to be absent.
Coded duties and work products
| Dimension | Exact category meaning | Postings in selected 100 (n) |
|---|---|---|
| Duty | Build or maintain intake or integration from named source systems. | 39 |
| Duty | Build, change, or operate executable data pipelines and jobs. | 79 |
| Duty | Design, build, or maintain analytical, logical, physical, dimensional, or metric models. | 48 |
| Duty | Implement or run explicit tests, validation, or reconciliation. | 35 |
| Duty | Operate monitoring or respond to data-system health, freshness, alerts, failures, or recovery. | 40 |
| Duty | Perform explicit query, job, platform performance or cost optimization. | 40 |
| Duty | Deliver or document a usable interface, model, or data product to a named consumer. | 5 |
| Output | A produced or maintained executable ingestion, transformation, or processing job. | 81 |
| Output | A curated dataset or data product expressly prepared for analytics consumption. | 25 |
| Output | A concrete model, schema, warehouse/mart model, or dimensional structure. | 47 |
| Output | Governed metric definition, semantic view, or dashboard-ready analytical layer. | 15 |
| Output | An explicit quality test, validation result, reconciliation, or control artifact. | 16 |
| Output | An explicit alert, health signal, monitoring system, or SLA-bearing output. | 13 |
| Output | A named technical document, lineage record, playbook, or runbook. | 30 |
The narrow explicit handoff category requires both a named recipient and an actual delivery, access or knowledge-transfer action; a product built for an unspecified purpose does not qualify. Quality testing means an assigned test, validation or reconciliation action; a general quality framework is coded separately.
What the postings disclose
A dimension is disclosed when the posting states source-codeable information in that field. Unstated means the accessible page does not say; it is not a negative finding. Unclear is kept separate. These are field-level disclosure counts, not proof that every specific category in the field could be assessed.
| Field | Disclosed | Unstated | Unclear |
|---|---|---|---|
| Education | 45 | 55 | 0 |
| Experience | 92 | 8 | 0 |
| Required criteria | 98 | 2 | 0 |
| Preferred criteria | 68 | 32 | 0 |
| Work cadence | 21 | 78 | 1 |
| Interfaces | 96 | 4 | 0 |
| Decision authority or escalation | 60 | 40 | 0 |
Paired mentions
For each pair, left and right are the numbers of selected postings supporting those separate categories; both is their same-posting intersection. The final column reports disclosure of both parent fields, not category-specific assessability. No pair rate is calculated, and a shared posting does not show that one practice causes the other. Pairs overlap with one another and with the rows above.
| Pair of coded categories | Left n | Right n | Both n | Both parent fields disclosed |
|---|---|---|---|---|
| Source ingestion × Analytics-ready dataset | 39 | 25 | 10 | 100 |
| Data modeling × Semantic or metric layer | 48 | 15 | 12 | 100 |
| Quality testing × Monitoring or alert output | 35 | 13 | 6 | 100 |
| Technical requirement translation × Semantic or metric layer | 5 | 15 | 0 | 100 |
| Governance, access or lineage work × Mapping or contract | 38 | 14 | 9 | 100 |
| Operating monitoring or recovery × Incident communication | 40 | 2 | 1 | 96 |
| Pipeline construction × Engineering change practice | 79 | 55 | 43 | 100 |
The engineering boundary of an analytics role
The common boundary in these advertisements is the handoff from operational or external data to a maintained analytical asset. A role can be close to business reporting and still be engineering-led when it builds production pipelines or reusable models. Taskrabbit's Analytics Engineer posting, for example, asks for dbt models that turn raw data into tested, documented data and for governed metric definitions shared with analysts and business stakeholders. The New York Times connects reusable business logic, data contracts, alerting, and reporting datasets to its analytics engineering position. These are materially different from a role limited to interpreting an existing dashboard.
At the platform end of the range, Amazon's DBS BI Data Engineer II is assigned warehouse and lake infrastructure, ETL from commercial and operational systems, a unified analytical model, and leadership dashboards. Walmart's Staff Data Engineer is asked to develop analytical pipelines and architectures, implement integration and quality frameworks, and guide data-modeling practice. These examples show why the source-to-consumption chain matters more than the title alone. They do not imply that every sampled role owns all of those stages.
Responsibilities and work products
Bringing source data into a usable system
Several postings make the starting point explicit: an engineer must identify, ingest, and reconcile data from systems whose structures and update patterns differ. HexArmor's Data Engineer I supports ingestion and transformation pipelines, works with business teams on requirements, and documents data flows. Its qualifications include relational SQL foundations, ETL familiarity, and MuleSoft integration experience. The duties and the candidate requirements are related but distinct: a named integration tool in a qualification is not, by itself, the output of the job.
Splice describes modernization of legacy jobs and production models supporting finance and revenue data. Fora Financial asks its staff engineer to reason about APIs, change-data capture, backfills, idempotency, schema drift, and failure recovery while designing and operating a platform. Those details show that ingestion is not simply copying rows. Source changes, retries, backfills, and the ability to recover from failure determine whether downstream analysis remains trustworthy.
Other settings add domain-specific integration constraints. Waymark's Senior/Principal Data Engineer must design systems resilient to inconsistent and delayed clinical data across electronic health records and payer sources, using healthcare interface and terminology standards where the role specifies them. Amazon Music Finance describes SQL, PySpark, and Airflow pipelines that turn varied inputs into structured datasets for finance analysis. These are examples of source integration under different operating conditions; their domain requirements should not be projected onto unrelated employers.
Integration can also be an agreement about meaning and safe movement across teams. Fusion Worldwide describes data flowing through a modeled object layer and back into a system of record; its requirements explicitly name transactional integrity, idempotency, reconciliation, and APIs used by applications the engineer does not control. Accenture Federal Services asks for source identification and deduplication, mapping, and API or SDK integration in a classified environment. 9amHealth connects pipelines and analytical models to clinical, product, finance, and engineering questions. The work product in such cases is a maintained path whose consumers can understand what was moved, how records were reconciled, and which definitions they may rely on. That interpretation follows from the linked duties; it does not assume a shared architecture among the employers.
Models that preserve business meaning
The postings repeatedly distinguish raw storage from a model that a business user can interpret. Amazon Music Finance explicitly calls for a consolidated data model so metric logic has one maintained source. Amazon Studios Analytics assigns pipelines and logical and physical models for content-performance datasets, with title-level launch measures and consistent upstream metric definitions. Walmart's Sunnyvale Senior Data Engineer connects analytical models to dashboard-ready tables and documents the architecture, flows, and transformations.
Analytics engineering examples make this semantic work especially visible. Digible asks for modeled layers and a governed metrics layer while assigning BI-tool strategy and enablement. OnePay assigns production dbt models, Databricks dashboards, and semantic metrics intended for self-service. Taskrabbit likewise links tested dbt models to a shared definition of measures. In these cases, a useful output is not merely a table: it is a reusable representation of entities and measures that other people can query without reconstructing the underlying logic.
Quality, observability, and operational ownership
An analytical dataset can be available yet wrong, late, or difficult to trace. The selected examples treat quality as work performed during construction and operation. Amazon Studios Analytics assigns ongoing dataset-quality monitoring and service expectations for reporting surfaced in applications. Amazon Music Finance names KPI and ETL-job monitoring systems with freshness checks, anomaly detection, and lineage. The New York Times includes data contracts and alerting alongside its modeled reporting data.
The operational burden is explicit in some postings. Splice names an on-call rotation and incident triage. Gridware also describes participation in on-call rotations. Judi Health asks for pipeline-health monitoring, incident troubleshooting, and root-cause fixes. These sources support event-driven incident work for those roles. They do not establish an on-call norm across the full sample, because many other postings do not disclose such a cadence.
The output should be read precisely. A posting that says “run quality checks” establishes a duty, while one that names implemented tests, monitoring systems, or documentation supports a more specific work product. OnePay assigns ownership of quality through testing, CI/CD, and documentation, but that wording alone does not name a separately delivered document. HexArmor, by contrast, explicitly asks the engineer to document processes and flows.
Skills, technologies, and qualifications
SQL, Python, modeling, transformation, orchestration, testing, and cloud data infrastructure appear in many of the linked examples, but a list of product names can be misleading unless its grammar is preserved. The study codes two things separately: the section containing the name and the status of the individual product within that section. A required qualification can ask for a category while offering several interchangeable examples. Splice, for example, asks for experience with a workflow orchestrator and names Dagster, Airflow, Prefect, or an equivalent. The requirement is the orchestrator capability; the posting does not make each example individually mandatory.
Other postings make the distinction particularly clear. 540 asks for Databricks or a similar modern data platform. Ondo Finance calls BigQuery strongly preferred within a broader cloud-warehouse requirement. Waymark expressly marks Docker as required while listing Kubernetes as an example in the same qualification. Walmart's Sunnyvale role describes cloud-platform experience and marks Google Cloud Platform preferred inside its required-skills sentence. Such details change the interpretation of a tool table; a product's presence in a posting is not an estimate of mandatory hiring demand.
The hard-skill crosswalk records 72 of the selected 100 postings as naming Python as a programming capability or an acceptable one-of option. This is a named-capability mention count, not a count of postings that individually require Python. Anduril Industries and OneSix illustrate alternative-language wording. An individual mandatory requirement needs its own source and qualification-status check. Generic labels such as “qualification” and “years of experience” remain linked to their individual records rather than being promoted into thematic frequency counts.
The distinction also applies to levels and credentials. HexArmor describes an entry-level Data Engineer I and asks for foundational SQL and limited relevant experience. Waymark separates senior and principal experience thresholds and adds principal expectations about technical direction and high-risk integration leadership. Accenture Federal Services specifies alternative education-and-experience paths for a classified-environment role. Those statements belong to their particular openings. The study does not turn a senior-heavy sample into a universal entry requirement, and it does not treat a preferred degree or certification as mandatory.
Seniority in a title should therefore be read with the actual decision and delivery language. Oscar Health's Analytics Engineer I asks for experience with SQL, semantic modeling, transformation tools, and cloud database tools while identifying additional healthcare and dimensional-modeling experience as bonus qualifications. Waymark adds principal-specific expectations to a shared senior/principal posting. Walmart's Staff role includes technical guidance on modeling, storage, and metadata practice. These examples illustrate distinct scopes—hands-on implementation, ownership, and technical direction—without making an employer's title ladder interchangeable with another's. Formal eligibility requirements, such as U.S. citizenship or active TS/SCI clearance on particular federal postings, must likewise stay attached to those openings rather than being recast as general data-engineering skills.
Interfaces, judgment, and work rhythm
Business analytics data engineering is often described as work across organizational boundaries. Amazon DBS BI names product, finance, service engineering, and sales partners whose analytical requirements the engineer supports. Gopuff's Data Engineer connects analytics, product, engineering, and operations to curated datasets. OnePay asks engineers to turn messy business problems into durable data products with data and analytics colleagues. The observable skill is not a vague label such as “good communication”; it is the ability to clarify a decision need, choose a representation, explain a trade-off, and make the resulting data usable.
The postings do not grant every engineer the same authority. Babylist explicitly says the senior individual contributor will own the architecture and systems that scale data engineering work. Fusion Worldwide describes an engineer who scopes a feature, builds a proof of concept, iterates, and decides when it ships in coordination with a product manager. Amazon Studios Analytics uses single-threaded ownership language for content-performance data and participation in modeling trade-offs. By contrast, a posting that merely says “design a model” establishes a duty; it does not automatically establish approval authority over a team or policy.
Work rhythms are similarly specific. OnePay explicitly assigns daily use of AI tools in development. NinjaTrader names end-of-day partner reporting. Splice states on-call participation. A candidate qualification saying someone already uses a tool daily is not automatically an assigned daily work schedule. Where cadence is unstated, the study records that limit rather than filling it from industry convention.
What this evidence can guide
For an employer describing such a position, the postings suggest the value of naming the expected artifact and its consumer. “Build pipelines” is less informative than specifying whether the engineer will maintain ingestion from changing source systems, publish a modeled analytical dataset, define governed metrics, or operate freshness and recovery controls. The Amazon Studios and OnePay postings illustrate that specificity in different settings. A clear role description can also separate a required capability, a preferred product, and an illustrative item from the current stack. That separation makes the hiring signal more interpretable without narrowing the candidate pool by accident.
For a practitioner comparing openings, the same discipline helps reveal what must actually be demonstrated. An engineer may need to show a reliable ingestion design, an explainable model, documented quality checks, or collaboration across product and finance—depending on the posting. Fora Financial emphasizes platform judgment and recovery; Digible emphasizes metric governance and BI enablement; Fusion Worldwide emphasizes end-to-end feature scoping and auditable data. This is an interpretation of stated work products and responsibilities, not a new requirement attributed to every employer.
Interpretation and limitations
This is a purposive reading of public hiring language, not a probability sample, a census of U.S. openings, or a forecast. Search visibility, applicant-tracking indexing, posting turnover, explicit geography, and the date of capture affect which vacancies could be selected. The senior-title tilt and the contribution of multiple requisitions from some employers shape the evidence. One Home Depot posting carried an older embedded target-end date, while its employer page stated ongoing applications and its exact Workday route still opened the application start screen on 5 October 2026; the study retains that source with this currentness limitation. Title variants are useful for identifying related work but cannot establish standardized occupational levels. A job advertisement states what an employer seeks; it is not direct observation of a person's daily work.
The coding also has deliberate limits. A posting can combine several responsibilities, omit a routine activity, place a preference inside a generally required section, or describe a current technology stack without requiring prior experience in every component. Required, preferred, illustrative, alternative, and current-stack mentions therefore remain separate. Some product lists use wording such as “for example” or “or similar” without making any one name individually required or preferred. We retain those names as illustrative, alternative, or unresolved mentions according to the posting wording and exclude them from individual requirement counts. Duties, outputs, and skills can overlap within one posting and must not be added as if they were mutually exclusive categories. Unknown fields remain unknown. The linked examples support the interpretations attached to them; they are not evidence that every employer uses the same tools, governance model, escalation route, or delivery rhythm.
This report draws its principal claims from the selected vacancies. Separate research on current vendor releases and technology changes answers a different question and does not validate the prevalence of a practice in U.S. vacancies. Employer and vendor names here identify sources, not endorsements. We use short evidence anchors and links, and do not reproduce full advertisements or protected training materials.
Conclusion
The sampled postings describe a role that engineers the path from heterogeneous sources to dependable analytical use. Its tangible products can include pipelines, modeled datasets, governed measures, monitoring, documentation, and self-service interfaces, depending on the opening. Its professional judgment appears in choices about source fit, data meaning, quality, performance, and handoffs with business and technical partners. The most defensible interpretation keeps each of those expectations attached to the employer statement that supports it and keeps silence, alternatives, and preferences visible.