Nexto Labs — agents for the work nobody wants
Hand over the job, not the tool.
We build agents that own one recurring job end to end — read the queue, decide, act, and escalate when they’re unsure. For ops and engineering leads at 50–500 person companies, where the job scales in headcount.
it follows the cursor. it stays when you scroll.
Is your job agent-shaped?
Four things have to be true at once. Most jobs a company would like to automate fail on the third one, and they fail quietly — the work looks routine until you try to write down what a correct outcome is.
Read the four below against a real job. If two of them are false, we’ll tell you on the call rather than three months into a build.
- High volume
- It happens dozens of times a week, not four times a month.
- Judgment-light
- Two experienced people would agree on the right outcome.
- A clear right answer
- You can grade it afterwards without a meeting.
- Reading and typing
- Someone does it today by looking things up and writing them down.
One agent, in full: inbound support triage.
There’s no client name on this one. We’re new, and a logo we don’t have wouldn’t tell you anything — so here is the mechanism instead, at the level of detail that would be uncomfortable to fake.
The job. Someone opens the support queue every morning and reads every new ticket. Most are one of eight things: where is my order, change the shipping address, the code didn’t apply, cancel this, refund this, the file won’t download, how do I do X — and something genuinely new. Sorting them, pulling the order, and answering the first seven is reading and typing. The eighth needs a person, and the cost of the job is that the person doing the eighth spends most of the day on the other seven.
How it’s checked. The eval set exists before the agent does: real tickets from the last quarter, each with the outcome the team agreed was correct — right intent, right action, right tone. Every change is scored against it. A build that drops below the previous score doesn’t ship. The set is versioned in your repository and it is yours when we leave.
When it’s wrong. Three failure classes, three answers. Wrong intent routes to a human and the ticket joins the eval set the same day. Wrong action inside a right intent can’t get far, because the actions that can’t be undone are proposals a person approves. A confidently wrong answer to a customer is the one that costs — so the agent cites the passage it answered from, uncited answers are blocked, and every send is sampled and read for the first weeks it’s live.
| Ticket queue | Read, search, tag, assign. | Cannot delete or merge. |
|---|---|---|
| Order lookup | Read-only. | No writes, ever. |
| Refund API | Drafts a refund. | A person approves every one. |
| Shipping address | One write, pre-dispatch. | Blocked after dispatch. |
| Knowledge base | Reads and cites. | An uncited answer is blocked. |
| Reply | Sends on the seven known intents. | The eighth goes to a human. |
Nothing in the right-hand column is a setting. Each one is a boundary the agent has no tool to cross.
- Inbound support triageRead the queue, classify, answer the known, escalate the rest.
- Prospect researchTake a list of companies, return the same twelve fields, with sources.
- Document intakePull structured fields out of PDFs that arrive in forty layouts.
- Internal question answeringAnswer from the wiki, cite it, and say “I don’t know” out loud.
The jobs we take on.
Four archetypes, because they’re the ones where the standard for “done correctly” can be written down. If yours isn’t here and it passes the four tests above, it’s still worth an hour.
How we work.
Six stages. The order matters more than any of them individually — stages two and three are the ones everyone wants to skip, and they’re the ones that decide whether anybody can tell if this worked.
Write the job spec
Describe the job the way you’d brief a new hire — including how you’d know they had done it badly.
Baseline the work as it is today
Volume, hours, cost, error rate. Skip this and there is no way to prove later whether the agent won.
Build the eval set before the agent
A graded set of real cases with known-correct outcomes. Clients underestimate this one and end up depending on it most.
Build against the evals
The agent, its tools, its guardrails — scored every iteration, not at the end.
Shadow run
The agent works the real queue alongside the person doing it now. Output is compared; nothing reaches a customer. It ends when it beats the baseline.
Graduated handover
Autonomy raised in steps, with escalation paths, a kill switch, and monitoring for drift. Then we operate it, or your team does.
What it costs.
Three shapes, in order. You can stop after the first — a lot of engagements should.
| Engagement | Price | What you get |
|---|---|---|
| Job spec + eval set2–3 weeks | to confirm | You leave with the spec, the measured baseline, and a graded eval set that is yours. If the job turns out not to be agent-shaped, this is the stage where we say so. |
| Build + shadow run6–10 weeks | to confirm | The agent, its tools, its guardrails, and the shadow run against your baseline. It graduates when it beats the person doing the job today, or it doesn’t graduate. |
| OperateMonthly | to confirm | We run it, watch for drift, and keep the eval set current as the job changes. Or we hand all three over and your team does. |
Model tokens and infrastructure are billed to you by your own providers. We don’t resell either.
When is an agent the wrong tool for this?
What happens the first time it gets something wrong?
Who owns the code, the prompts, and the eval set afterwards?
Does it run in our cloud or yours, and what do the model providers see?
What does it cost to run per month, not just to build?
How much of our team’s time does the build consume?
The awkward questions.
Answered here so you don’t have to spend the call extracting them. The first one is the one we care most about getting right.
Book a working session.
An hour. Bring one job you’d hand over. You leave knowing whether it’s agent-shaped, what its eval set would have to contain, and the first place it would go wrong. If it isn’t a fit, that’s the answer you get.
or hello@nxtolabs.com — answered within one business day.