B

Senior Data & LLM Engineer (AI‑Ready Data & Agentic Systems)

Blue Orange Digital · Anywhere

Full-timeLeadPythonAWSGCPAzureDockerOpenAI

About the Role

Company overview:Blue Orange Digital is a boutique data & AI consultancy that delivers enterprise-grade results. We design and build modern data platforms, analytics, and ML/AI Agent solutions for mid‑market and enterprise clients across Private Equity, Financial Services, Healthcare, and Retail. Our teams work with technologies like Databricks, Snowflake, dbt, and the broader Microsoft ecosystem to turn messy, real-world data into trustworthy, actionable insight. We're a builder‑led, client‑first culture that prizes ownership, clear communication, and shipping high‑impact work. Note: Please submit your resume in English, as all application materials must be in English for review and consideration. Position overviewBlue Orange Digital is seeking a Senior Data Engineer with genuine agentic systems depth for a full-time, 6-month embedded engagement in a private-capital data environment. This is a deliberately merged role: roughly half senior data engineering — Snowflake, dbt, Airflow, AWS, Python ingestion — and half senior LLM/agentic engineering, spanning production agents, MCP servers, evaluation practice, and observability. You will work on two core initiatives: an AI-Ready Data Catalog & Context Layer built alongside the Data Analytics team, and Unstructured Data Extraction at scale, with agentic implementation, evaluation, and orchestration as the throughline across both. This is not a pipeline-maintenance seat; we are looking for someone who has already shipped agents into production and can tell us candidly what broke. ResponsibilitiesAgentic development (primary focus)Build production agents and multi-step agentic workflows, and design and implement MCP (Model Context Protocol) servers and tool interfaces as the primary integration surface between agents and the data platform — including custom internal MCP servers over proprietary datasets. Make deliberate architectural calls on agent vs. deterministic pipeline, single-agent vs. multi-agent decomposition, and where human-in-the-loop checkpoints belong. Agent evaluation and observabilityStand up an evaluation practice from scratch — eval datasets, LLM-as-judge, regression testing — and define and track the metrics that matter: task success rate, faithfulness, extraction accuracy, tool-call correctness, latency, and cost per run. Instrument tracing and observability across agent runs. Agentic orchestrationOwn agentic workloads in Airflow 3.x on Astronomer — modeling LLM and agent calls as named, independently retriable tasks, using dynamic task mapping for fan-out/fan-in, assets and data-aware scheduling for event-driven triggers, deferrable operators for long-running calls, and human-in-the-loop operators for approvals. Design the workloads that do not belong in a DAG (long-running or interactive agents on AWS ECS/Fargate, queue- and event-driven execution, API-triggered services) and articulate clearly where the orchestration boundary sits and why. AI-ready catalog and semantic context layerModel and implement semantic definitions over a Snowflake + dbt estate — metric definitions, entity relationships, business glossary, and ownership — so agents receive governed meaning rather than raw table names. Build and maintain Snowflake Semantic Views and Cortex Analyst semantic models, keeping them in sync with dbt as the source of truth, and establish how context is versioned, tested, and promoted. Metadata and governance instrumentationCapture lineage, freshness, column-level descriptions, classification and sensitivity tags, and usage signals; automate description and tag generation with LLMs where it is safe to do so, behind human review gates. Evaluate catalog and metadata tooling and deliver a clear build-vs-buy recommendation. Unstructured data extractionDesign extraction pipelines for PDFs, decks, spreadsheets, emails, and web content, including layout-heavy financial and operational documents where tables and nested structure matter. Build schema-constrained extraction using structured outputs and tool calling with frontier LLM APIs and Cortex LLM functions, with confidence scoring, citation back to source page, and explicit handling of low-confidence fields. Implement chunking, embedding, and retrieval strategies serving both RAG and agentic retrieval, and design the human review loop — what auto-accepts, what escalates, and how corrections flow back into eval datasets. Senior data engineering and architectureDefine and implement architectural solutions that ensure the delivery speed and scalability of data pipelines, including tiered/medallion structures; drive best practices for advanced ELT, data governance, and data quality; and define and enforce data security and access-control policies across pipelines and platforms. Build and maintain Python ingestion pipelines (dlt and similar) across REST APIs, SaaS connectors, file drops, and incremental loads into Snowflake, and contribute to dbt models, tests, and documentation — treating

Blue Orange Digital has 1 open position on Remote Vibe Coding Jobs.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →