Nexto Labs — agents for the work nobody wants

Hand over the job, not the tool.

We build agents that own one recurring job end to end — read the queue, decide, act, and escalate when they’re unsure. For ops and engineering leads at 50–500 person companies, where the job scales in headcount.

Book a working session One hour · one job · a straight answer

it follows the cursor. it stays when you scroll.

Is your job agent-shaped?

Four things have to be true at once. Most jobs a company would like to automate fail on the third one, and they fail quietly — the work looks routine until you try to write down what a correct outcome is.

Read the four below against a real job. If two of them are false, we’ll tell you on the call rather than three months into a build.

High volume
It happens dozens of times a week, not four times a month.
Judgment-light
Two experienced people would agree on the right outcome.
A clear right answer
You can grade it afterwards without a meeting.
Reading and typing
Someone does it today by looking things up and writing them down.

One agent, in full: inbound support triage.

There’s no client name on this one. We’re new, and a logo we don’t have wouldn’t tell you anything — so here is the mechanism instead, at the level of detail that would be uncomfortable to fake.

The job. Someone opens the support queue every morning and reads every new ticket. Most are one of eight things: where is my order, change the shipping address, the code didn’t apply, cancel this, refund this, the file won’t download, how do I do X — and something genuinely new. Sorting them, pulling the order, and answering the first seven is reading and typing. The eighth needs a person, and the cost of the job is that the person doing the eighth spends most of the day on the other seven.

How it’s checked. The eval set exists before the agent does: real tickets from the last quarter, each with the outcome the team agreed was correct — right intent, right action, right tone. Every change is scored against it. A build that drops below the previous score doesn’t ship. The set is versioned in your repository and it is yours when we leave.

When it’s wrong. Three failure classes, three answers. Wrong intent routes to a human and the ticket joins the eval set the same day. Wrong action inside a right intent can’t get far, because the actions that can’t be undone are proposals a person approves. A confidently wrong answer to a customer is the one that costs — so the agent cites the passage it answered from, uncited answers are blocked, and every send is sampled and read for the first weeks it’s live.

What it’s given, and where it stops
Ticket queueRead, search, tag, assign.Cannot delete or merge.
Order lookupRead-only.No writes, ever.
Refund APIDrafts a refund.A person approves every one.
Shipping addressOne write, pre-dispatch.Blocked after dispatch.
Knowledge baseReads and cites.An uncited answer is blocked.
ReplySends on the seven known intents.The eighth goes to a human.

Nothing in the right-hand column is a setting. Each one is a boundary the agent has no tool to cross.

  • Inbound support triageRead the queue, classify, answer the known, escalate the rest.
  • Prospect researchTake a list of companies, return the same twelve fields, with sources.
  • Document intakePull structured fields out of PDFs that arrive in forty layouts.
  • Internal question answeringAnswer from the wiki, cite it, and say “I don’t know” out loud.

The jobs we take on.

Four archetypes, because they’re the ones where the standard for “done correctly” can be written down. If yours isn’t here and it passes the four tests above, it’s still worth an hour.

How we work.

Six stages. The order matters more than any of them individually — stages two and three are the ones everyone wants to skip, and they’re the ones that decide whether anybody can tell if this worked.

  1. Write the job spec

    Describe the job the way you’d brief a new hire — including how you’d know they had done it badly.

  2. Baseline the work as it is today

    Volume, hours, cost, error rate. Skip this and there is no way to prove later whether the agent won.

  3. Build the eval set before the agent

    A graded set of real cases with known-correct outcomes. Clients underestimate this one and end up depending on it most.

  4. Build against the evals

    The agent, its tools, its guardrails — scored every iteration, not at the end.

  5. Shadow run

    The agent works the real queue alongside the person doing it now. Output is compared; nothing reaches a customer. It ends when it beats the baseline.

  6. Graduated handover

    Autonomy raised in steps, with escalation paths, a kill switch, and monitoring for drift. Then we operate it, or your team does.

What it costs.

Three shapes, in order. You can stop after the first — a lot of engagements should.

EngagementPriceWhat you get
Job spec + eval set2–3 weeksto confirmYou leave with the spec, the measured baseline, and a graded eval set that is yours. If the job turns out not to be agent-shaped, this is the stage where we say so.
Build + shadow run6–10 weeksto confirmThe agent, its tools, its guardrails, and the shadow run against your baseline. It graduates when it beats the person doing the job today, or it doesn’t graduate.
OperateMonthlyto confirmWe run it, watch for drift, and keep the eval set current as the job changes. Or we hand all three over and your team does.

Model tokens and infrastructure are billed to you by your own providers. We don’t resell either.

When is an agent the wrong tool for this?
When the volume is low, when the right answer is genuinely contested, when the inputs were never written down, or when being wrong once is unrecoverable. A job that happens four times a month never pays back its eval set. A job where two experienced people disagree about the correct outcome has no gradeable standard, so there is nothing to build against. We would rather lose the deal here than in month three.
What happens the first time it gets something wrong?
You hear it from monitoring, not from a customer — because it ran in shadow first, and because the actions that can’t be walked back are proposals a person approves. The wrong case goes into the eval set the same day, which means that specific mistake is now something every future build is scored against.
Who owns the code, the prompts, and the eval set afterwards?
You do. All three, in your repository, from the first commit. The eval set is the part that matters: it is the only artefact that keeps its value if you replace the model, the framework, or us.
Does it run in our cloud or yours, and what do the model providers see?
Yours. The agent is deployed into your cloud account, so the data path stays inside your perimeter apart from the model call itself. What reaches a provider is whatever is in that call, under that provider’s commercial terms, and not used for training. If a class of data must never leave, we scope the agent so it never reads it.
What does it cost to run per month, not just to build?
Model tokens plus whatever it is deployed on, billed to you by your own providers — we don’t resell either. We estimate it from your baseline volume during the spec stage, before you commit to the build, and we try to be wrong in the direction of over-estimating.
How much of our team’s time does the build consume?
Concentrated at the start. The spec and the eval set need a few hours a week from the person who does the job today — nobody else can supply the correct answers. After that it settles into a review cadence.

The awkward questions.

Answered here so you don’t have to spend the call extracting them. The first one is the one we care most about getting right.

Book a working session.

An hour. Bring one job you’d hand over. You leave knowing whether it’s agent-shaped, what its eval set would have to contain, and the first place it would go wrong. If it isn’t a fit, that’s the answer you get.

Book a working session

or hello@nxtolabs.com — answered within one business day.