You build the agent loops that write most of the code and content at this company, and you own whether a loop's output is safe to merge. The truth verdicts on shopping claims belong to a sister seat; the loops belong to you.

Product.ai is the verified truth layer for shopping, the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes, its first proof at scale, is the code verification service, at around $22M in revenue. The product is codes that actually work, proven by robots running real checkouts. Founder-owned and profitable since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.

Why This Role Exists

Agents write most of the code and most of the content here. That only works because a loop layer decides what those agents can reach, and a gate layer decides whether what they produce is allowed to ship. Nobody owns that layer full time yet. You do.

You build the orchestration that lets an agent plan a task, call the right tool, read back what happened, and try again when it's wrong. You build the retrieval layer that lets an agent answer from the company's own knowledge instead of guessing. And you build the eval gates that stand between a loop's output and production, so speed never quietly becomes sloppiness. This is not a research seat: every system you build runs today, on real work that ships.

This role is not "prompt the model and ship the reply." A prompt that works once and breaks on the next input is not a system. You are building the infrastructure other engineers and agents build on, so your failures compound and your fixes compound too.

The System You'll Need to Model

  • Agent orchestration and tool use. A loop that plans a multi-step task, calls a tool, reads the result, and decides what to do next differs from a single prompt-response call. It needs state, retries, timeouts, and a way to know when to stop and ask for help. You design that control flow, not just the prompts inside it.
  • Retrieval-augmented generation over a live, growing brain. Cortex, the shared AI brain the company runs on, answers its own questions from more than 8,600 internal documents and grows daily. Your retrieval layer has to stay accurate as the corpus changes: what to index, how to chunk it, when an answer is grounded versus invented.
  • Prompt engineering and AI systems architecture as a discipline, not a craft trick. A retrieval chunk and an MCP tool result fill an agent's context window differently, and each fails differently when the budget blows: a dropped chunk means the agent guesses instead of citing, a truncated tool result means it acts on partial information without knowing it. You design what goes in, what gets summarized, and what gets dropped, and you can explain why an upstream change altered what the agent did downstream.
  • Eval-gated shipping. Agents write much of the code and content here; your corpora and gates decide whether a loop's output ships or gets rejected. You build beside the seat that owns the verdict science behind those gates; you build the loops that produce the work those gates judge.
  • MCP servers that AI shopping agents call. Other people's AI agents reach into this company's verified data through the Model Context Protocol. You build and harden the servers on the other end of that call, including what happens when a caller sends something malformed or hostile.
  • A company that changes its own rules while you work. Since August 2026, any operator here can propose a change to how Cortex works and ship it live, without founder approval. Your systems have to keep working while the rules under them move, so you design for change instead of a stable spec.
  • A high bar for reasoning before code. You will work beside the company's two lead architects, whose standard is computer-science depth: reasoning about correctness, invariants, and failure modes before a line gets written, not after.

If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.

What You Will Own

  • The agent loop architecture. The orchestration layer that plans, calls tools, and recovers from failure across the systems that write this company's code and content. You decide how loops are structured, retried, and stopped, and you own the incidents when a loop does something it shouldn't.
  • The retrieval layer. What an agent can know and how fast it can know it. You own chunking, indexing, freshness, and the judgment call on when an answer should say "I don't know" instead of guessing.
  • The eval gates on loop output. Working beside the seat that owns verdict science, you build the harnesses and regression corpora that decide whether a given loop's output is safe to ship. Unlike the deterministic build-gates on the infrastructure side, these run a judge, because whether a loop's output is right is subtler than whether a test assertion passed. A gate that never fails anything is not a gate.
  • The MCP servers. The interface other AI agents use to call this company's verified data. You own its correctness, its failure modes under bad input, and its behavior under load as more of the shopping web calls it.
  • Your seat charter. Within your first quarter you co-sign a charter for this seat: one machine-checkable number that proves it's working, and a written split of what you decide freely versus what you bring to the founder.

You will use the craft you already own (LLM application engineering, TypeScript and Python, API design) and grow into the layer above: deciding when a gate's judgment should be overruled by a human versus by another model, and extending the MCP surface to the next protocol as more of the shopping web starts calling in.

Who You Are

Handed an agent loop you didn't write, you can trace what a tool call actually did versus what it claimed to do, and you notice fast when the two disagree. You write clearly, because the next person debugging your loop at 2am will be reading your reasoning, not just your code.

You verify a retrieval chunk against its source before trusting the citation, and you verify a tool call actually did what it reported before trusting the loop's next step; an agent's word is never enough for either. You can build what an agent is supposed to build, by hand, which is exactly what lets you tell whether what it handed back is right. You move between architecture and shipped code in the same day: a retrieval design in the morning can be a deployed change by evening. Compute is cheap here. A redo cycle from an unverified loop is not.

You have probably built an agent or tool-use system that ran in production and not just a demo, a retrieval pipeline that stayed accurate as its source data changed, or an API that other systems, human or automated, depended on without babysitting. Adjacent roads count: chatbot platforms, developer tooling, search infrastructure, or any system where you had to reason about what a model would do with unseen information. We care about the artifact and the reasoning more than where you built it.

Who this isn't for. This is wrong if your instinct is to ship a working demo and call the hard part done, because here the demo is the easy 20% and the gate that makes it safe in production is the job. It's wrong if you want a single well-scoped lane and a ticket queue, because this seat spans orchestration, retrieval, and evaluation, and "that's not my system" ends the conversation. It's wrong if you'd trust a retrieval answer because it sounds grounded, or trust a tool call succeeded because the loop didn't throw; both need checking against the source and the result, not the confident tone. It's wrong if "I built an agent framework" is your credential rather than a loop that survived a caller sending it something malformed. And it's wrong if you need the spec fully written before you can start, because the rules here change while you're building. You'll be happiest here if you'd rather build the guardrail than argue about whether one is needed.

How We Evaluate

We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.

  1. Async video screen. About 15 minutes, on your own time. It replaces the recruiter screen. We want to see how you think, not how you present.
  2. Calls with company stakeholders. Short conversations with the people you'd build beside.
  3. Conversation with the founder. How you reason about the systems above, and whether you can hold the argument live.
  4. Paid work trial. Real work in our real environment: a live loop, live retrieval, or a live eval gate, not a take-home puzzle. We watch how you get grounded fast, whether you write the spec before the build, and whether your self-assessment is honest.

If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.

Compensation & Ownership

Total first-year comp: $400,000 to $500,000 (base plus performance-based ownership and profit-share programs). Base: $250,000 to $320,000, top of market for senior AI engineering.

Eligibility for the company's ownership and profit-share programs. Grants are performance-based, terms discussed at the offer stage. 100% premium coverage for you and your family. Your token budget is effectively unlimited, steered by return, never capped.

This structure is built to mint partners. Based in Santa Monica, Los Angeles, in person, five days a week. Relocation support available for the right builder.