W

Senior Software Engineer - AI Innovation

Worth AI · Anywhere

🔥9 people viewed this job

About the Role

Worth AI, a leader in AI onboarding and underwriting, is looking for a talented and experienced Senior Software Engineer - AI Innovation to join our team. At Worth AI, we are on a mission to revolutionize decision-making with the power of artificial intelligence helping fintechs, lenders, payment processors, and financial institutions onboard small businesses faster, smarter, and more confidently. We're building the infrastructure that powers real-time KYB, KYC/IDV, underwriting, and continuous risk monitoring at enterprise scale, and the Worth Score™ our unified credit score derived from 1,200+ data points across 700M+ SMBs. As a Senior Software Engineer - AI Innovation, you won't be wiring up demos you'll be designing and shipping production agent systems that make consequential decisions on regulated financial data. Worth's platform consolidates onboarding and underwriting into a single AI-powered system, and our agents read, reason over, and act on the messy, high-stakes signals that come with that domain. You'll own the end-to-end lifecycle: architecting agent graphs, building the retrieval and tool layers they rely on, instrumenting them with evals and observability, and getting them deployed against SOC 2 / GDPR / CCPA guardrails. You'll partner closely with our Chief AI Officer, applied scientists, product, and platform teams to turn agentic patterns into customer outcomes. Responsibilities Design and ship multi-step agentic systems (planner/executor, tool-using, multi-agent, human-in-the-loop) that automate KYB, underwriting, case review, and risk monitoring workflows.Architect agent graphs in LangGraph (or comparable frameworks CrewAI, AutoGen, Claude Agent SDK) with explicit state, durable execution, retries, and safe fallbacks.Build and harden the retrieval layer powering our agents chunking strategies, hybrid search, reranking, and grounded citation across SoS filings, IRS records, bank data, and Worth's 700M+ SMB graph.Own the eval stack: golden sets, offline regression suites, LLM-as-judge, online A/B and shadow evals, and red-teaming for jailbreaks, prompt injection, and PII leakage.Wire agents into Worth's production systems via well-typed tools, MCP servers, and existing services (decisioning engine, case management, crosswalking). Treat tool surface area as a product.Drive production MLOps for agents: deployment, versioning, traffic shaping, cost/latency budgets, observability (traces, token spend, tool call success), and on-call playbooks for agent incidents.Partner with security, compliance, and legal to keep agents inside Worth's SOC 2, GDPR, CCPA, and fair-lending posture — building from day one, not bolted on.Translate ambiguous product bets ("what if the underwriter had an AI co-pilot for this?") into concrete agent designs, prototypes, and shipped features.Mentor engineers across the org on agent patterns, prompt engineering hygiene, eval discipline, and the failure modes of LLM systems.Stay ahead of the frontier new models, frameworks, and patterns — and bring back what actually works in production.Technology Stack Languages & Runtimes: Python, Node.js, TypeScriptAgent / LLM frameworks: LangGraph, LangChain, Claude Agent SDK, MCP, OpenAI SDKModels: Anthropic Claude, OpenAI, open-weight (Llama, Mistral) where appropriateRetrieval & Data: PostgreSQL, pgvector / vector DBs, OpenSearch, Kafka, Redshift, RedisInfra & Orchestration: AWS, Kubernetes (EKS), ArgoCD, TerraformEvals & Observability: LangSmith / Langfuse / Braintrust-style tooling, DataDog, custom eval harnessesRequirements 8+ years of professional software engineering experience, with at least 2 years building production LLM or agentic systems (not just notebooks or demos).Solid software engineering experience - front-end, APIs, async patterns, queues, databases, and the failure modes of distributed systems.Demonstrated ownership of major features or subsystems in production.Demonstrated experience mentoring junior engineers and raising team quality standards.Demonstrated experience with event-driven systems: enrichment, retries, dead-lettering, backpressure.Experience managing containerized applications in Kubernetes, EKS, ArgoCD, operators, Kustomize.Deep, hands-on experience with at least one modern agent framework (LangGraph strongly preferred) and a track record of shipping agents that actually run, fail gracefully, and recover.Real experience with evals you've built golden sets, run offline and online evaluations, and used them to make ship/no-ship calls.Production MLOps fluency: you've deployed LLM workloads under real latency, cost, and reliability constraints, and you instrument what you ship.Strong proficiency in Python; comfortable in TypeScript / Node.js for integrating with Worth's services.Clear, calibrated communicator - able to explain agent trade-offs to product, security, and customers without hand-waving.Operates with extreme ownership in ambiguous, fast-moving environments. Excited to work alongside a

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

⏰

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →