Skip to content
Matheus Prates

Time in Goiânia: Goiânia

PTRésumé
Back to the projects

02Tenaz

No customer left hanging.

An app that reads the WhatsApp conversations of people who sell and keeps a list of what they promised, what's expected of them and what no longer matters.

The AI only reads; it never replies for you. Three rounds a day with Claude, a queue in PostgreSQL and old conversations archived as Parquet.

  • React Native
  • Expo
  • Express
  • Graphile Worker
  • Claude
Commitments screen in dark mode: what is due to be received and paid, and a question on whether an old commitment still stands
Conversations screen: the test account's customers, with what is still unread and the stage of each deal
Tenaz Today screen: two new conversations to approve, a question about a stalled commitment and five people waiting for an answer

The problem

People who sell over WhatsApp promise to get back to someone all day long. Among hundreds of messages, one forgotten promise becomes a customer who bought from the competition.

What I built

  • An Android and iOS app built with React Native and Expo.
  • Connection to the person's WhatsApp through Evolution API. Messages show up in the app instantly, without going through AI.
  • Three rounds a day, at 8am, noon and 7pm, where Claude reads only what's new, together with the account's memory.
  • Every commitment comes with a quote from the conversation, and the API checks that the quote exists before saving it.
  • When a matter is closed, the commitments tied to it are reviewed and closed automatically.
  • Audio transcribed by Whisper; conversations older than 180 days move to Parquet in cheap storage.

The screens

Screenshots of the real system. Click a screen to see it full size.

  • People screen: 14 people in the pipeline, split by stage, with search and filters

    People: the pipeline by stage, from new contact to closed deal.

  • Results screen: five commitments kept and a chart of kept commitments per month

    Results: what got cleared and how much of what was promised was delivered.

  • What the AI reads screen, in dark mode: eight of thirteen conversations switched on for reading

    What the AI reads: the user picks every conversation that goes into the reading.

  • Connection screen: WhatsApp connected, 13 conversations, 683 messages and 12 days of history

    Connection: the WhatsApp status and what is already stored.

  • Pairing screen in dark mode: the eight-letter code to type into WhatsApp

    Pairing by code, on the same phone, with no second device needed.

How it works

  1. A message arrives

    The webhook writes it to PostgreSQL and the app shows it right away.

  2. Audio becomes text

    Whisper, running on Groq, transcribes voice notes.

  3. The round

    Three times a day, one Claude call per account, with what hasn't been read, the memory and the state of each conversation.

  4. The check

    The API checks every quote against the right conversation and only then saves commitments, topics and decisions.

Decisions that mattered

  1. The AI only reads

    Why: Replying to customers on someone's behalf is a risk nobody wants to take. The app reminds you; you do the replying.

  2. Three rounds a day

    Why: Tested on 14 days of real conversations: the same cost as one daily read, with three times as many reads and the average wait dropping from 9 hours to 2.9.

  3. Niche: people who sell

    Why: Measured on real conversations, the app pays off most for people who handle lots of customers on WhatsApp. New features only go in if they help them.

What went wrong

Every real system breaks somehow. These were the stumbles that taught the most, and what changed because of them.

  1. The made-up test said 88%. The real one, 54%

    What happened

    I tuned a small model's prompt by looking at conversations I had written myself. On real conversations the same reading got far less right, and half the error came from a rule I had written just to pass the test.

    What changed

    I only tune with real conversations exported from production. And I learned that, with a small model, adding rules made things worse four times in a row; what helped was removing and reordering.

  2. The queue stopped without an error

    What happened

    A daily queue cleanup ran while the worker was up and deleted the task types it had cached at startup. New tasks sat still, with no error and no lock, and the reading rounds didn't run for a day and a half.

    What changed

    Type cleanup only runs when the queue starts. And I learned to watch for absence: a round that never happens is a failure too, even with no error message.

  3. The local model didn't pass the benchmark

    What happened

    I tried running the reading on an open model, on a Mac at home, to bring the cost to zero. On the sales-stage step it scored 60.7% against Claude Haiku's 82.1%, and it made up sale totals nobody had mentioned.

    What changed

    The reading stayed on Claude. No model goes into the product without first passing the benchmark of real conversations.

Where it stands

Beta in production. The Play Store release comes next.