Product.ai is the verification layer for commerce, and our shopping chat is where a shopper asks a product question and gets an answer with the evidence behind it. Our test team rates the answers. Each rating saves a record with the question the tester asked, the rating itself, and any note they wrote.

A nightly classifier, a language model, reads each new record that has a question or a note and files it under 1 of 7 jobs, such as “Decide what to buy” or “Protect against a bad outcome,” or marks it as fitting none. We wrote those 7 families in June 2026, when the record store was new. On 2026-08-18, the first time we grouped the labeled records into themes, 57 of the 104 records in the run were in “Decide what to buy,” and as of 2026-09-16 that family has 135 of the 240 labeled records. No person has confirmed a single one of the classifier’s labels.

You read every record yourself and decide what each tester was trying to get done. Then you tell us whether those 7 families are the right ones.

Our goals for this mission

The chat has to feel like talking to a friend who knows products, and that starts with knowing what people come to it for. The job families are how we count that. They fill a grid on our feedback dashboard, and the themes built on top of them are how fixes get routed.

Part of the reason so many records fall into “Decide what to buy” is a measurement problem. On 2026-08-18, only 14 of 214 records had a tag for what went wrong. The pass that splits each family into smaller themes had nothing to work with, so we now fill that gap with a machine-written tag. What a machine cannot tell us is whether “Decide what to buy” is a single job or several, and whether the labels are right. That takes a person who reads the records.

The records so far come from our own team. Outside testers join in small, hand-approved groups, and their feedback goes into the same store. We want the families right before their questions arrive in volume.

What you would do

  • Access and setup. You get read access to the record store and the classifier’s output, and you read the 7 family definitions and the 3 theme files from June to August 2026. You ask the product manager who runs the chat what decisions the families feed.
  • The read. You read every record and write the job each tester was trying to get done in your own words before you look at the machine’s label. Then you mark each label right or wrong and note where you cannot tell.
  • The revised families. You group your own job statements into a revised set of families and name the jobs the current 7 miss or merge. The records the classifier marked as fitting no family are the first place to look.
  • The re-run. You run the revised families back over the records and report how the distribution changes.
  • The handoff. You write the handoff. It has the revised families with a one-line definition each, your agreement rate with the classifier, and a trace from every family to the records behind it.

What success looks like

  • A revised set of job families exists, each with a one-line definition anyone could apply to a new record without asking you.
  • Every family traces to real records, and every record you moved out of the biggest family has a reason next to it that anyone could check.
  • You report your agreement rate with the classifier, with the disagreements listed, and you say plainly if the 7 families survive your read and the problem is elsewhere.
  • The product manager who runs the chat can read the handoff in one sitting and decide what to change.

Ground rules

  • The deliverable is the revised families with their traces, not a research report. Scope, timing, and pay are set in the contractor agreement, at your stated rate.
  • You write your own job statement for each record before you see the machine’s label. Labels graded after the fact do not count.
  • The records are stored in a private repository. Getting read access is the first task, and if it is refused you say so at once.
  • Notes from our own team are attributable by name. For records from outside testers, you attach a name to a note only where that tester opted in, and you never quote or attribute a question an outside tester typed.
  • You may use AI tools to help you read and sort, and every family you propose still traces to records you read yourself.
  • The evidence is the records, so interviews and surveys are out of scope.

What happens after you send a proposal

3 steps, and a person reads every proposal.

  1. You send a short written proposal. It covers how you understand this mission and how you would approach it, and you attach your resume or a LinkedIn profile.
  2. We read it and reach out with questions. There is no test and no exercise.
  3. If there’s a fit, the mission begins under a contractor agreement and an NDA, paid at your stated rate, with kickoff in Santa Monica.