# RAG and Enterprise Search Work in 2026: A U.S. Vacancy Study

> A study of 105 current U.S. employer vacancies from 67 employers shows how RAG, enterprise search and knowledge-system work is divided across engineering, research and delivery.

- Canonical page: https://mtfinstitute.com/insights/rag-enterprise-search-us-vacancy-study-2026/
- Content type: Article
- Editorial category: Research &amp; Reports
- Publisher: MTF Institute of Management, Technology and Finance
- Author: MTF Institute Research Team- Published: 2026-10-04
- Updated: 2026-10-05
- Language: English
- Topics: Retrieval-Augmented Generation, Enterprise Search, Knowledge Systems, Information Retrieval, U.S. Vacancy Research

**Author:** MTF Institute Research Team  
**Independent review:** MTF Institute Research QA  
**Evidence date and geography:** 4 October 2026, United States  
**Technical report number:** MTF-CF-RR-2026-10-04-RAG-SEARCH  
**Research design:** Structured purposive vacancy study

**Version 2 correction (4 October 2026):** Four accepted postings offer a non-U.S. location or eligibility alongside a U.S. option. The combined sensitivity subset contains 81 requisitions, rather than the 82 shown in version 1. The 105-record source cohort and the substantive conclusions are unchanged.

## Abstract

What does work on retrieval-augmented generation (RAG), enterprise search and knowledge systems look like in current U.S. job advertisements? We examined 105 distinct, U.S.-eligible employer requisitions from 67 employer clusters that were publicly accessible on 4 October 2026. Each accepted advertisement stated material retrieval, search or knowledge-system work in the role, rather than merely naming RAG in a preferred-skills list. The corpus is a purposive, cross-sectional sample, not a census or a national estimate.

The advertisements describe several related kinds of work: preparing and connecting source information, designing search and retrieval behavior, combining retrieval with applications and assistants, evaluating answer or search quality, and operating systems within security, cost and performance constraints. They also describe different professional positions in that work. Some engineers own an index or service; researchers investigate retrieval or relevance methods; data and knowledge specialists prepare usable source structures; client-facing specialists adapt methods to customer settings; product and architecture roles make broader design decisions. The visible outputs include pipelines, services, search features, evaluation arrangements, interfaces and design decisions. These are not interchangeable job titles or a single uniform occupation.

The study&#039;s strongest conclusion is therefore a description of connected work, not a claim that one technology, tool or credential dominates the U.S. labor market. It identifies practical evidence a professional should be able to produce: a defined information problem, a traceable source and access model, a retrieval design appropriate to that problem, quality tests, and a clear account of operational trade-offs. Employer wording and selected examples support this conclusion; individual advertisements often leave cadence, authority or handoff details unstated.

## Research question and scope

The research question was: **What duties, methods, outputs, working relationships and decision boundaries do current U.S. employer vacancies explicitly associate with RAG, enterprise search and knowledge systems?** We chose the United States as the study geography before analysis. A vacancy was eligible when the employer&#039;s direct page showed a current application path, a U.S. location or explicit U.S. eligibility, and a material role duty involving retrieval-augmented generation, search or indexing, retrieval-ready source preparation, knowledge representation, relevance or grounding evaluation, or authorization-aware retrieval.

This definition reaches beyond titles containing “RAG.” An associate scientist building graph-backed retrieval and an engineer designing a search index may both contribute to knowledge-system work, but their responsibilities and authority differ. Conversely, a broad AI role was not eligible merely because its team used RAG or its qualifications mentioned a vector database. We retained adjacent public-web search and customer-facing solution roles when the particular vacancy stated transferable retrieval work, and labelled their different operating setting. Their inclusion cannot establish how many internal enterprise employers use any method.

The unit of analysis is a requisition, identified by employer and provider job ID. The 105 accepted requisitions came from direct employer pages and employer-managed applicant-tracking pages. Fifty-two were in the Greenhouse/Ashby source slice and 53 in other employer-board slices. These are source-route counts, not two independent estimates of demand. The 105 records represent 67 employer clusters because several employers had multiple distinct requisitions. Four clusters—Amazon (10), ServiceNow (10), Exa (6) and LTS (5)—contributed 31 of the 105 records. Multiple openings from one employer may share language, staffing priorities and technology, so we report employer concentration alongside requisition counts.

## Sampling and coding method

Searches used RAG, retrieval-augmented generation, enterprise search, knowledge engineering, vector and hybrid search, retrieval evaluation, document ingestion and related terms. Candidate employer pages were opened and checked; search snippets alone were not accepted as evidence. The researchers excluded duplicate job IDs and mirrored URLs, closed openings, ambiguous U.S. eligibility, generic team descriptions and qualification-only references to retrieval. For each accepted page, the source record retained its employer, role title, provider job ID, URL, retrieval date, U.S. location or eligibility evidence and a short necessary excerpt. The analysis relies on paraphrase and limited excerpts, rather than reproducing advertisements.

Coding distinguished what the posting said the worker would **do**, what the worker would **produce**, what technical skills or products it named, and what human work or decision rights it described. Separate fields captured duties, outputs, hard skills, observable behavioral skills, tools, education and experience, required versus preferred status, seniority, cadence, interfaces and handoffs, and authority or escalation. An item was recorded as stated only when a specific employer-page passage supported it. “Unstated” means that the reviewed page did not support that field; it is neither evidence that the work never occurs nor a zero in an employer&#039;s practice. Required and preferred labels follow an explicit heading or wording on the employer page. A named product in “nice to have” is not a mandatory job requirement, and a product named in qualifications does not automatically become a duty.

The study also separates three levels of evidence that are easy to conflate. First, the accepted source cohort establishes that each of 105 roles contains material retrieval or knowledge-system work. Second, field coding records what the reviewed pages explicitly state about work and conditions. Third, narrower technical themes require direct support for that particular technique. The general retrieval/RAG inclusion screen is not an independent discovery about how common retrieval is. Counts of narrower themes, where used, must be read as support within retained employer-page excerpts, with the 105 accepted requisitions as the sample denominator; they cannot estimate prevalence across all U.S. vacancies. Absence from an excerpt is not absence from a complete employer role, still less absence from actual work.

We checked whether the picture changes when the largest employer clusters are removed and when contract-award-contingent work, cross-border eligibility, adjacent client-facing delivery and public-web search roles are treated separately. Those categories can overlap. Excluding the four largest employer clusters leaves 74 requisitions. A narrower view that removes the four labelled strata leaves 81. Neither subset is a representative enterprise sample; each tests whether an interpretation depends heavily on a visible concentration or on a different type of role. The distinctions are particularly important for claims about internal enterprise knowledge work.

The field-status audit indicates how much of the reviewed advertising text was available for each kind of question. It counts records with at least one supported entry in a field, **not** how many employers use a particular skill or practice. For example, a behavioral entry may describe collaboration in one role and mentoring in another; 81 cannot be read as a frequency of either behavior. Each row has the same 105-requisition denominator.

| Reviewed field | Requisitions with a stated entry | Unstated in the reviewed page |
| --- | ---: | ---: |
| Duties | 105 | 0 |
| Outputs | 105 | 0 |
| Hard skills | 102 | 3 |
| Observable behavioral skills | 81 | 24 |
| Tools or systems | 91 | 14 |
| Required versus preferred classification | 68 | 37 |
| Task cadence | 21 | 84 |
| Named interfaces or handoffs | 73 | 32 |
| Authority or escalation | 58 | 47 |

The uneven coverage explains why it is possible to describe visible work products yet inappropriate to assign a standard cadence or decision right to every role. A stated output field still needs a separate check before a particular item can be called a pipeline, design document or other deliverable. The table reports field availability, not validated thematic categories.

## The shape of the observed roles

An exclusive classification based on **title wording** puts 53 of 105 requisitions in software, AI, platform or search engineering; 24 in research or machine learning; 14 in client-facing solution delivery; seven in data or knowledge engineering; five in architecture or strategy; one in security engineering; and one in product management. These labels describe the observed titles and do not validate occupational boundaries. Work can cross the labels: an ML engineer may operate a retrieval service, and a data engineer may build a search feature. A practitioner planning work should therefore inspect the stated task and expected output instead of relying on the job title alone.

Title seniority is similarly diverse. Fifty-three titles lack an explicit seniority marker; 26 contain senior, eight staff, five senior staff, five principal, five manager or lead, and three associate, student or trainee markers. These groups sum to the 105 requisitions, but a marker is not a comparable measure of experience across employers. In particular, “senior” and “staff” do not establish people-management or approval authority. The role body, not the title, must support a conclusion about who makes architecture decisions, approves a release, assigns work or escalates an incident.

The range is visible in individual employer pages. [AbbVie’s associate scientist role](https://jobs.smartrecruiters.com/AbbVie/3743990014656470-associate-scientist-data-ii) (VAC-001) connects Cypher queries, data loading and entity resolution to graph-backed downstream retrieval. [Glean’s search-quality ML role](https://job-boards.greenhouse.io/gleanwork/jobs/4738120005) (VAC-037) discusses combining language models with search and training a model that uses ranking signals. [Glean’s API and context-platform role](https://job-boards.greenhouse.io/gleanwork/jobs/4696752005) (VAC-038) makes search and knowledge-graph capabilities accessible through API endpoints and states ownership of API standards. [Pi Security’s search and knowledge-systems role](https://jobs.ashbyhq.com/pi-security/c527b018-5a76-4157-9e82-0a9312a3d681) (VAC-068) connects chunking, hybrid search, reranking, metadata filters and citations to a knowledge system that represents domain entities and provenance. These are four distinct job designs, not four copies of an occupational template.

## From source material to retrieval behavior

The first recurring work boundary is between a source collection and something a user can reliably retrieve. Job advertisements describe pipelines that ingest and prepare information, indexes that make it searchable, and application paths that turn retrieval results into user-facing answers. That boundary matters because a response cannot be evaluated meaningfully without knowing which material was available, how it was transformed, which user was entitled to it, and what the retrieval step actually returned.

Several postings make this chain concrete. [Amazon’s Quick Suite software-engineering role](https://www.amazon.jobs/en/jobs/10557546/software-development-engineer-amazon-quick-suite) (VAC-015) describes scalable ingestion pipelines and retrieval optimization, as well as collaboration with applied scientists. [ServiceNow’s agentic-search infrastructure role](https://careers.servicenow.com/jobs/744000130592199/agentic-search-infrastructure-engineer-moveworks/) (VAC-078) describes infrastructure spanning an index engine, ingestion and enrichment. [Tiger Analytics’ senior data-engineering role](https://apply.workable.com/tiger-analytics/j/80C8ED5E26/) (VAC-097) describes data pipelines supporting language-model applications, vector embeddings and knowledge retrieval. These examples show that preparation and retrieval are connected tasks, but they do not imply the same data source, architecture, cadence or level of responsibility in each employer.

Retrieval design is not one technique. [Amazon’s OpenSearch Vector Search role](https://www.amazon.jobs/en/jobs/10523004/software-development-engineer-amazon-opensearch-vectorsearch-opensearch) (VAC-014) explicitly names high-performance vector-search algorithms and large-scale datasets. [Amazon’s DynamoDB index search and storage role](https://www.amazon.jobs/en/jobs/10561388/software-development-engineer-dynamodb-index-search-and-storage) (VAC-016) explicitly describes combining vector similarity, lexical relevance and structured attributes in hybrid retrieval and ranking. [Zip’s senior data-engineering role](https://job-boards.greenhouse.io/zipcolimited/jobs/4696276006) (VAC-102) names vector search, hybrid search and RAG capabilities in an enterprise data and AI platform. Each example supports a particular technique in that advertisement. It would be unsound to transfer the technique to every vacancy in the sample simply because all 105 met the broader retrieval-work inclusion rule.

The associated output may be a pipeline, index, service, API, search experience, knowledge model or design decision. [Glean’s API platform advertisement](https://job-boards.greenhouse.io/gleanwork/jobs/4696752005) (VAC-038) is an example of an integration surface as a deliverable. [AbbVie’s graph-focused role](https://jobs.smartrecruiters.com/AbbVie/3743990014656470-associate-scientist-data-ii) (VAC-001) emphasizes queries, validation and entity resolution. A report on this work should name the produced object and its users rather than treating “built a RAG system” as sufficient evidence of performance.

## Quality, security and operating constraints

Quality appears in more than one form. Search relevance, retrieval quality, grounded answers, hallucination checks and application performance call for different tests. [Exa’s evaluation-research role](https://jobs.ashbyhq.com/exa/9c45a74e-d507-482a-bc0e-da2f464c9767) (VAC-028) describes evaluation frameworks that probe search and an evaluation stack that supports feedback to research, infrastructure and product teams. Because Exa operates in public-web search, this is evidence about an adjacent search setting, not direct evidence of internal enterprise knowledge practice. [LTS’s AI platform and harness role](https://job-boards.greenhouse.io/lts/jobs/4337509009) (VAC-051) explicitly names retrieval quality, hallucination detection, latency, throughput, cost and application performance in evaluation work. [ServiceNow’s benchmarking and evaluation management role](https://careers.servicenow.com/jobs/744000137502479/senior-engineering-manager-agentic-generative-ai-benchmarking-and-evaluations/) (VAC-081) discusses testing grounding and RAG pipelines with search and data-fabric teams. These postings support a practical distinction between an attractive demo and a measured system, but the corpus does not establish any universal pass threshold or testing standard.

Security and permission boundaries are also visible in role wording. [Accellor’s AI principal-engineer role](https://apply.workable.com/accellor/j/B29FA60C33/) (VAC-002) names permission-aware retrieval as an architecture pattern and connects technical trade-offs with latency, reliability, safety and cost. [LTS’s agentic AI security role](https://job-boards.greenhouse.io/lts/jobs/4340457009) (VAC-050) explicitly concerns securing RAG pipelines, embeddings, vector databases and knowledge repositories; its stated monitoring output concerns abnormal agent behavior and manipulation. [Pi Security’s knowledge-systems role](https://jobs.ashbyhq.com/pi-security/c527b018-5a76-4157-9e82-0a9312a3d681) (VAC-068) refers to metadata filtering, grounding, citations and provenance. These are different forms of control: which document may be retrieved, how a result supports an answer, and how abnormal behavior is noticed. The vacancy descriptions do not establish that any one employer has already implemented all of those controls or that a named design is sufficient for security.

Operational constraints require the same caution. One [Amazon software-engineering advertisement](https://www.amazon.jobs/en/jobs/10557546/software-development-engineer-amazon-quick-suite) (VAC-015) offers an illustrative day that includes reviewing system metrics and performance dashboards. [LTS’s platform role](https://job-boards.greenhouse.io/lts/jobs/4337509009) (VAC-051) describes troubleshooting production AI issues and continuing reliability improvements. These examples support the existence of monitoring and improvement work in those roles. They do not justify assigning a daily dashboard review, on-call rota or release cadence to every other vacancy. In the coding record, most postings leave a specific work cadence unstated; that is a transparency limit, not evidence that operational routines are absent.

## Collaboration, authority and qualification signals

The visible work also moves across professional interfaces. The [AbbVie advertisement](https://jobs.smartrecruiters.com/AbbVie/3743990014656470-associate-scientist-data-ii) (VAC-001) describes partnering with analysts and domain experts to gather requirements. [Glean’s search-quality role](https://job-boards.greenhouse.io/gleanwork/jobs/4738120005) (VAC-037) includes mentoring more junior engineers. [Pi Security’s role](https://jobs.ashbyhq.com/pi-security/c527b018-5a76-4157-9e82-0a9312a3d681) (VAC-068) connects customer needs with product, backend, frontend, platform and security teams. [Exa’s evaluation role](https://jobs.ashbyhq.com/exa/9c45a74e-d507-482a-bc0e-da2f464c9767) (VAC-028) names ML researchers, data engineers, infrastructure engineers and product colleagues. These are observable interfaces, not a generic personality requirement. In each setting, useful evidence of competence would show how a requirement, test result or design trade-off was communicated to the people who must act on it.

Authority is less uniform. [Glean’s API role](https://job-boards.greenhouse.io/gleanwork/jobs/4696752005) (VAC-038) specifically assigns ownership of REST API standards and versioning decisions. [Accellor’s principal role](https://apply.workable.com/accellor/j/B29FA60C33/) (VAC-002) assigns architecture-pattern and trade-off ownership across several technical constraints. [Amazon’s DynamoDB search role](https://www.amazon.jobs/en/jobs/10561388/software-development-engineer-dynamodb-index-search-and-storage) (VAC-016) describes helping determine how indexes are created, maintained and recovered. These statements support distinct kinds and scopes of technical decision. A title alone cannot extend any of them to budget authority, staffing decisions or final approval of an organization-wide release.

Qualifications require equally literal reading. [Amazon’s Quick Suite role](https://www.amazon.jobs/en/jobs/10557546/software-development-engineer-amazon-quick-suite) (VAC-015) places a stated experience threshold under basic qualifications and a separate life-cycle experience statement under preferred qualifications. [LTS’s platform role](https://job-boards.greenhouse.io/lts/jobs/4337509009) (VAC-051) puts named vector-database products and orchestration frameworks in a “Nice to Have” section. Those employer-specific distinctions matter more than a list of popular tools assembled from all postings. The study cannot infer that every person in this work needs the same vendor stack, degree or years of experience. A qualification may be required in one requisition, preferred in another and unmentioned in a third.

## Concentration and sensitivity

The sample deliberately captures a range of retrieval-related roles, but it is concentrated enough to distort an unqualified summary. Amazon and ServiceNow each account for 10 of 105 requisitions; Exa and LTS add another 11. Together the four employers contribute 31 observations from four of the 67 clusters. A feature repeated across openings from one of these employers is not 10 independent confirmations that the feature is common across U.S. employers. Removing those four clusters produces a 74-requisition sensitivity set. The remaining roles still span information preparation, search or retrieval design, application integration, evaluation and collaboration, but a theme&#039;s exact count and emphasis must be recalculated in that set before any comparative claim.

The type of employer and role also changes the interpretation. Ten records belong to public-web search employers, where crawling, indexing and relevance work may be unusually central. Fourteen have adjacent client-facing delivery titles, where the professional may diagnose a customer&#039;s needs, integrate a solution or explain trade-offs rather than own a single internal retrieval service. One opening is contract-award contingent, one other has a contract title, and one is part-time. Four postings list a non-U.S. option alongside U.S. eligibility: a [Firecrawl machine-learning role](https://jobs.ashbyhq.com/firecrawl/72f9dc1d-65db-48c9-b3d9-c6ccdb997006) and [search role](https://jobs.ashbyhq.com/firecrawl/762b4426-b4aa-4377-96d3-51f40c59cbf7) name Toronto; [Liatrio](https://jobs.lever.co/liatrio/dabbe592-d22e-49ef-91a9-b775a3296fdd) names Canada; and [Lifted](https://jobs.smartrecruiters.com/LiftedanUpworkCompany/3743990014736041--132210-ai-platform-technical-lead) also permits parts of Latin America. These strata overlap, so their counts must not be subtracted as if they were mutually exclusive. The combined sensitivity view excludes contract-award-contingent, cross-border, adjacent client-facing and public-web-search strata and contains 81 requisitions; the separate contract-title and part-time markers are reported without being automatically excluded. That subset helps test claims specifically about internal U.S. enterprise roles; it does not turn the purposive sample into a representative survey.

The evidence supports a family resemblance among the roles rather than one canonical “RAG engineer” profile. Source preparation, retrieval behavior, answer quality, access and operations form a connected system, but employers split responsibility across functions. A vacancy study can show which combinations were **stated** in these current requisitions. It cannot show the proportion of all U.S. employers that use a particular architecture, whether one product will prevail, or whether a learner will obtain a job.

## Interpretation and practical implications

The findings suggest five practical questions for anyone scoping this work. **What is the information problem?** Define the users, source collection, update need and decision the search or answer experience supports. **What is retrievable?** Specify ingestion, representation, index and permissions so that a result can be traced to an authorized source. **How is relevance or answer support judged?** Use a test set and observable criteria suited to the task, including failure cases rather than one persuasive example. **What must the system connect to?** Identify interfaces, documentation and handoffs needed by data, product, security or customer-facing colleagues. **What are the operating constraints?** State latency, scale, cost, monitoring and recovery requirements where the actual workplace task calls for them.

These questions are derived from the kinds of duties and outputs visible in the sample. They are not a prescribed vendor implementation or a claim that every vacancy includes every step. An entry or mid-level contributor might be asked to prepare source data, implement a bounded retrieval component and record test results. A senior technical owner may need to set architecture boundaries, explain trade-offs and coordinate a release. A research specialist may develop relevance or evaluation methods. Client-facing work may put more weight on requirements discovery and integration. The appropriate evidence of performance changes with the role, while a traceable relation between source, retrieval behavior, test and decision remains useful across them.

The broader U.S. occupational context is consistent with the need to connect software work to user needs, design, testing, maintenance and communication. [O*NET&#039;s Software Developers profile](https://www.onetonline.org/link/summary/15-1252.00) and [detailed work context](https://www.onetonline.org/link/details/15-1252.00) describe those activities for a broad occupation; the [U.S. Bureau of Labor Statistics profile](https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm) offers another broad description of software development work. Neither source identifies a separate RAG occupation or measures the frequency of RAG work. We use them only to interpret why vacancy language about testing, collaboration and maintenance matters, not to fill an unstated field in any of the 105 records.

## Limitations

This is a point-in-time, purposive sample of visible employer advertisements, not a probability sample. Search terms, ranking, employer-board access and the time openings remain live influence what could be observed. A requisition describes an employer&#039;s advertised intent; it does not directly observe an employee&#039;s practice or prove that an organization has deployed a production system. Distinct employer job IDs avoid duplicate openings but do not remove correlation among posts from one company. Title grouping is a transparent organizing device, not a validated occupational classification. The cross-border, contract, public-web and client-facing strata call for separate interpretation.

The coding is bounded by employer wording. Many pages omit cadence, approval rights, detailed handoffs or whether a tool is mandatory. A short retained excerpt supports a positive observation but may not exhaust every relevant passage in a full advertisement. An unstated code must therefore stay unknown. Specific technical frequencies, if reported, need direct passage support for the precise activity or output: the word “retrieval” is not proof of hybrid search, “architect” is not automatically a design document, and a general security term is not automatically evidence of a concrete access-control deliverable. The report avoids making such conversions.

This study does not estimate national vacancies, employer adoption, hiring growth, wages, organizational return, technology efficacy or employment outcomes. It does not determine legal or regulatory requirements. Product releases and research published during 2026 can explain changing technical possibilities, but that separate evidence is not part of the vacancy numerator and cannot turn a feature announcement into a claim about U.S. employer prevalence. Any future comparison would require a declared sampling design, repeated observation and consistent coding.

## Research archive and data

The report and a rights-safe factual index of the 105 public employer source URLs are preserved in the [MTF Institute Zenodo record](https://doi.org/10.5281/zenodo.23144487). Read the [archival PDF report](https://zenodo.org/records/23144487/files/rag-enterprise-search-us-vacancy-study-2026.pdf?download=1) or [source URL index](https://zenodo.org/records/23144487/files/accepted-vacancy-index-rights-safe.csv?download=1). The index contains identifiers, titles, locations and links; it does not reproduce employer job descriptions. The version DOI identifies this dated evidence package, not a continuously updated count of open jobs.

## Conclusion

The 105 observed vacancies show retrieval and knowledge-system work distributed across engineering, research, data, architecture, product, security and solution-delivery settings. Specific advertisements connect source preparation, retrieval methods, search or assistant features, evaluation, access boundaries, operational constraints and cross-functional decisions. Their scopes vary, and several important work conditions are often unstated. The most defensible professional reading is to ask what information system a role must build or improve, what evidence would show that it works for authorized users, and who is responsible for the decisions around it. The sample supplies concrete examples of that work while leaving national prevalence and future demand unanswered.

## Continue learning

Apply the evidence from this report through MTF Institute&#039;s [Professional Certificate in RAG, Enterprise Search &amp; Knowledge Systems](https://mtfinstitute.com/programs/rag-enterprise-search-knowledge-systems/#enroll). The programme turns the identified capabilities into structured theory, guided AI practice and reusable workplace artifacts.



## Citation

When citing or summarizing this material, link to the canonical HTML page: https://mtfinstitute.com/insights/rag-enterprise-search-us-vacancy-study-2026/
