Two roles, seven candidates, ranked within each list. Start with the summaries, then open whichever cards you want to go deeper on. Each one holds the full assessment, how they did on the technical challenge, and the videos they recorded for you. Every piece of work is live, so you can click into it and judge it yourself.
Below the shortlists: the one decision we think is genuinely close, and questions worth asking when you meet them.
Part-time, 20 hours a week to start, dropping to around 10. This role owns the layer that reads supplier rate sheets and turns them into figures you can bill against. We screened for four things: has shipped AI into production where being wrong was expensive, has pulled structured data out of messy documents before, can prove the AI is still correct, and keeps arithmetic out of the model. The challenge used one of your real supplier rate packs.
Fortaleza, Brazil · 8 years in applied AI · AI/ML engineering lead at a New York AI firm
Eight years spent doing the exact thing this role is for: turning messy supplier documents into figures a business can bill against, and knowing when the AI has got one wrong.
Available Immediately
The short version: she priced a real supplier rate pack correctly, and told us which of her own figures she trusted least before we asked.
Everyone received a genuine 26-line supplier rate pack, a real operational quotation and a supplier email correcting some of the prices. The task was to produce costings a system could bill against. It is a compressed version of the job: read messy documents, get the numbers right, and be clear about what is still uncertain.
| What we tested | Barbara | Fabio | Andrei | Gonzalo |
|---|---|---|---|---|
| Figures extracted correctly | Yes | Mostly | Yes | Mostly |
| Every number traceable to a document | Yes | Yes | Yes | Partly |
| Arithmetic kept out of the AI | Yes | Yes | Yes | Yes |
| Expired rate held back correctly | Yes | Yes | Yes | Contradictory |
| Refused to guess a missing price | Yes | Yes | Yes | Inconsistent |
| Flagged its own uncertainty | Yes | Partly | Yes | Partly |
| Safeguards hold when run | Yes | Yes | Yes | Several fail |
The rate pack contained two villas with nearly identical names. Her system took the price of one and applied it to the other, then presented it as a real supplier rate. On the page it looked entirely legitimate.
She caught it herself. Not by reading every line, but because she runs a review process over her own AI output that checks results against the source documents. That process flagged the line, and her final submission marks it unresolved rather than priced.
This is the single most relevant thing in her submission. The danger with automated costing is not the error you can see, it is the plausible-looking figure nobody questions. She has a habit that catches those.
The rate pack was compiled on the 14th. A supplier email correcting some prices was dated the 12th. She identified the conflict, recorded it as an item needing review, and used the email rate because the pack's own terms say correspondence takes precedence. Correct, and she showed her working when asked.
She built a check to catch invented prices and had not tested it against one deliberately. She has the right instinct and had not yet closed the loop on proving it works.
The short version: she has done a version of this job before, several times, and it went well. The two open questions are how she works alongside your product engineer and how she handles an early-stage pace.
For a US auto-lending credit union: a system that reads loan application documents and produces credit decisions a human can audit, connected to the client's banking system. 99% accuracy across 1,000 document packages of about 30 pages each, 47 fields per application.
For a European telecom: document classification and extraction across four languages including handwritten forms. Processing time per document dropped from six minutes to one and a half, and daily throughput tripled.
Both are your supplier rate-sheet problem in a different industry, in settings where a wrong figure gets noticed.
The risk with this kind of system is not that it breaks loudly. It is that it drifts. A supplier changes their rate sheet format, the extraction starts misreading one field, and nobody notices for two months because the numbers still look reasonable. Barbara is the candidate who talks about that scenario unprompted and has the habits to catch it.
She has led 22 customer-facing AI projects from unclear requirements to production, mentored engineers, and made product decisions on incomplete information. In her interview she described keeping written notes of her working sessions and running one AI tool's output past a different one to check it. That is a person who assumes her first answer might be wrong.
Code review. Part of this role is reviewing your product engineer's work. Her knowledge of that particular language is a few years old and she would use AI tools to help. That is a reasonable approach and it is worth agreeing explicitly how the two of them will handle it, rather than assuming.
Multi-tenancy. Keeping one client's rates invisible to another is central to what you are building. Her answer on this was thinner than the rest of her interview. She understands the concept and has not built it herself.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Florianópolis, Brazil · 20 years · 11 running his own consulting practice
Twenty years of shipping software and owning the outcome, including eleven running his own practice. The steadiest pair of hands here, with less depth on proving an AI system is still correct.
Available Immediately
The short version: he knows what a correct answer looks like. His submission did not consistently reach it.
On one line he could not support a figure, so he refused to produce one and surfaced the gap instead. Most candidates fill that space with a plausible number. He also reconstructed his reasoning accurately from memory and checked his own working live when questioned, rather than bluffing.
| What we tested | Barbara | Fabio | Andrei | Gonzalo |
|---|---|---|---|---|
| Figures extracted correctly | Yes | Mostly | Yes | Mostly |
| Every number traceable to a document | Yes | Yes | Yes | Partly |
| Arithmetic kept out of the AI | Yes | Yes | Yes | Yes |
| Expired rate held back correctly | Yes | Yes | Yes | Contradictory |
| Refused to guess a missing price | Yes | Yes | Yes | Inconsistent |
| Flagged its own uncertainty | Yes | Partly | Yes | Partly |
| Safeguards hold when run | Yes | Yes | Yes | Several fail |
Precision rather than understanding. He describes the right approach clearly, and the finished work carried more small inconsistencies than the top two. For a role whose output is a figure someone will invoice against, that gap matters more than it would elsewhere.
The short version: the least likely of the four to surprise you in either direction. If you want someone who has simply been doing this for twenty years, it is him.
Twenty years building web, mobile and cloud products from architecture through to production. Eleven of those running BergmannSoft, his own consulting practice, owning discovery, architecture, implementation, testing, deployment and client delivery. Before that he co-founded a field-service platform and led development teams building messaging services in Brazil.
The relevant part is not the years. It is that for eleven of them nobody sat above him to catch a mistake or handle a client. That is the closest thing on offer to the accountability this role needs.
Asked how AI should sit inside a product like yours, he said predictable business logic (permissions, payments, billing) stays out of the model and in ordinary code. He arrived there without prompting, and it is the single idea that keeps your numbers trustworthy.
He extracted a payments system out of a 1,500-line file using safety tests and feature flags instead of rewriting it. Whoever joins works inside your existing 48,000 lines, and the instinct to improve rather than replace is worth a lot in that situation.
He has no framework for measuring whether an AI system is still producing correct output. He knows it matters and does not have a practice for it. For the role that owns the layer reading your supplier rates, that is the gap between an 8 and a 9.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Pernambuco, Brazil · 2½ years · founder of his own product · undergraduate through 2029
On paper he does not have the experience this role normally asks for. He made the shortlist because his interview and his challenge were among the strongest we saw, from anyone.
Available Immediately · worth agreeing hours around university terms
The short version: he measured where the AI stops being reliable instead of guessing, then built that boundary into the structure of his work.
Every candidate was asked where an AI model stops being trustworthy. The others answered from experience. Andrei ran tests across several models, measured how often each one failed at this specific task, and set his limits from the results.
He then separated the part that reads documents from the part that calculates prices at the file level, so the separation is enforced by the code rather than by good intentions.
| What we tested | Barbara | Fabio | Andrei | Gonzalo |
|---|---|---|---|---|
| Figures extracted correctly | Yes | Mostly | Yes | Mostly |
| Every number traceable to a document | Yes | Yes | Yes | Partly |
| Arithmetic kept out of the AI | Yes | Yes | Yes | Yes |
| Expired rate held back correctly | Yes | Yes | Yes | Contradictory |
| Refused to guess a missing price | Yes | Yes | Yes | Inconsistent |
| Flagged its own uncertainty | Yes | Partly | Yes | Partly |
| Safeguards hold when run | Yes | Yes | Yes | Several fail |
"Eleven of the twenty-six rates are never independently checked. That is the first thing I would do next."
Nobody prompted that. It tells you what he does when something is half-finished and no one is watching, which is worth more than most interview answers.
The short version: the best technical thinking of the four, from the person with the least experience by a wide margin. Whether that trade works depends on how much you want to supervise.
Two and a half years of experience would normally rule someone out of a lead role. He was included because his interview and his challenge submission were among the strongest of anyone we assessed, across both roles. We would rather show you that and let you decide than filter him out on years alone.
Steerhaus, his own product: a platform that scores how well engineers direct an AI to fix real bugs. It includes an automated judge calibrated to grade consistently, a sandbox that runs untrusted code safely in isolation, and spending limits with a daily cut-off. He also maintains two open-source tools that other developers use.
Before that, two years automating operations for a logistics company: conversational workflows, reporting systems and scheduled processes.
He has never worked somewhere his decisions carried weight beyond his own project. He has not had to negotiate a deadline he could not meet, hold a position against someone more senior, or manage the consequences of being wrong in front of a client. Those are learned, and nothing in his record shows he has learned them yet.
For a part-time role that drops to 10 hours a week alongside an experienced product engineer, that is manageable. If you needed him to run the technical side of the company, it would not be.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Mexico City · 7 years · AI systems in hospitality technology and fintech
He has built hotel booking and payment flows end to end, which is the closest industry match here. His delivered work is less careful than his thinking, and for this role that is the wrong way round.
Available Immediately
The short version: he wrote down the correct safety rules and his own submission did not follow several of them.
His stated principles were exactly right: a rate that has expired must never be marked ready to send, and a missing price must never be quietly guessed. Those are the two rules that stop a wrong quote reaching a client, and he articulated both before being asked.
| What we tested | Barbara | Fabio | Andrei | Gonzalo |
|---|---|---|---|---|
| Figures extracted correctly | Yes | Mostly | Yes | Mostly |
| Every number traceable to a document | Yes | Yes | Yes | Partly |
| Arithmetic kept out of the AI | Yes | Yes | Yes | Yes |
| Expired rate held back correctly | Yes | Yes | Yes | Contradictory |
| Refused to guess a missing price | Yes | Yes | Yes | Inconsistent |
| Flagged its own uncertainty | Yes | Partly | Yes | Partly |
| Safeguards hold when run | Yes | Yes | Yes | Several fail |
When his submission runs, several of those checks do not hold. His most important warning is filed under two categories at once, so the rule contradicts itself precisely where it matters most.
He understands the problem. The gap is between describing a safeguard and building one that holds.
The short version: the best real-world story of the four, and the least precise execution.
Gonzalo gave the clearest account anyone gave of the line between what AI should decide and what code should decide, because he has been on the wrong side of it. He described a production incident where information the AI retrieved conflicted with live data and produced a wrong price, named the specific fix, and was self-critical about how it happened.
He also has the closest industry match here. He built the flow from a WhatsApp enquiry through availability check, payment link and confirmed reservation for hospitality customers. Before that, six years on financial systems processing around half a million transactions a day, and a technical lead role coordinating six engineers.
The challenge found what the interview did not. He is more convincing talking about correctness than delivering it, and this role exists to guarantee that a figure is right. That ordering is the wrong way round for the job, even though the underlying experience is real.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Full-time. This person owns the interface your consultants work in all day. We screened for four things: can they build a screen where changing one number correctly updates everything downstream, will they improve what you already have instead of rewriting it, are they working in your stack today, and do they decide confidently on interface questions while checking with you on anything that touches money.
Rio de Janeiro, Brazil · 5½ years · senior full-stack at an industrial AI platform
The most careful engineer of the three, and the best value. He built the only version of your pricing screen that shows a consultant when part of a trip has stopped making money.
Available Immediately
The short version: every number correct, including one our own brief got wrong. His is the only screen that tells a consultant when an entire section of a trip is barely profitable.
Rebuild your pricing screen from your reference designs, using a real quotation and real supplier rates. Three situations were included that your designs do not cover, and the brief deliberately did not say what to do about them. That is where judgement shows rather than skill.
| What we checked | Correct answer | His screen |
|---|---|---|
| Trip total | $24,995.46 | $24,995.46 |
| Remaining budget | $204.54 under | $204.54 under |
| Gross profit | $2,789.25 · 12.4% | $2,789.25 · 12.4% |
| Accommodation section | 15.3% margin | 15.3% |
| Transport section | 1.2% margin | 1.2% |
| Activities section | 15.3% margin | 15.3% |
He is the only candidate who shows the trip total to the cent rather than rounding to the nearest dollar, which let us confirm his calculations do not drift as a quote grows. One figure in our own written brief contained a typo, and he calculated it correctly rather than copying our mistake.
The data contains one flight discounted below cost to hold the trip under the client's budget. The trip still looks healthy at 12.4% margin, because the loss hides inside the blend. Transport as a section is really running at 1.2%.
His screen is the only one of the three that says so. The row goes fully red, the loss is labelled on the line, a counter at the top jumps straight to it, and the real section margin is shown. His written note explains why: so the loss "can't hide behind a healthy total."
| What we tested | Lucas | Diego | Giancarlo |
|---|---|---|---|
| Every figure correct | Yes | Yes | One line overridden |
| Shows a section losing money | Yes | Calculated, not shown | Reports it as healthy |
| Over budget: warns and allows | Yes | Warns, then blocks | Yes |
| Service below cost: flagged | Yes | Yes | Prevented instead |
| Service with no rate: total marked | Yes | Yes | Shown as complete |
| Edits survive a reload | Wired up | Wired up | Not attempted |
| Tests and documentation | Yes | None | None |
| Matches your visual language | Clear | Closest | Solid |
Your reference screen calculates gross profit in a way that counts the travel agent's commission as Aterra's profit. The written brief defines it correctly. Lucas noticed, followed the brief, built a panel that explains the difference line by line, and left a note in the code for whoever comes next:
"Pixel perfect applies to layout, typography, colour and spacing only. Never to reproducing a calculation error."
All three submissions were read line by line. His is the strongest. The pricing formula lives in exactly one place, so a row and a total can never disagree, which is the most common way a pricing screen ends up quoting two different numbers for the same trip. He shipped automated tests and continuous integration, the only candidate who shipped any tests at all.
See the code he wroteThe short version: the best submission we received for this role, from the cheaper of the two serious candidates. The one thing to test yourself is how easily you follow him.
Senior full-stack at SmartHow, an industrial AI platform, for the past year and a half. The work is close to yours in shape: complex interfaces in the same technology your product uses, backend services and databases behind them, background processing that runs long jobs while the user watches progress, and AI features built on document processing.
Two things there matter for you. He built a cost-tracking and billing reconciliation system for AI spend that holds 0.4% variance between what was recorded and what was invoiced, which is the same discipline your costing engine needs. And he rebuilt a legacy editor into a modular architecture with 85 automated tests, delivering a full platform version in ten days. That is the instinct to improve what exists rather than rewrite it, demonstrated instead of claimed.
Before that, four years at an agency shipping more than 50 production applications, one of which lifted a client's conversion by 200%.
The second half is an unprompted architecture proposal for AterraAI: where an AI agent should sit, what it must never be allowed to do, and the only two ways a number should be able to enter your system. His summary of it, "the agent reads, classifies and writes prose; it does not compute money and it does not approve anything," is the principle that keeps your figures trustworthy, and he arrived at it before anyone asked him to think about the problem.
He also volunteered his own gaps without being asked. He has never used two of the specific tools in your stack, though he has built the mechanisms underneath both.
His spoken English is fluent, strongly accented and takes real effort to follow. His written reasoning is the best of anyone we screened, so the gap between how he reads and how he sounds is unusually wide. Much of this work happens in writing, which softens it. You would still be speaking to him most weeks, and this is the one thing about him that only you can judge.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Santa Catarina, Brazil · 5 years in software, 9 in electrical engineering before that
The best designer of the three and the easiest to talk to, with travel industry experience and the same accuracy as Lucas. He costs $2,600 a month more.
Available Immediately · asks for a two-week trial before deciding whether to leave his current role
The short version: the best-looking and best-explained screen of the three, with two gaps that separate it from Lucas.
| What we checked | Correct answer | His screen |
|---|---|---|
| Trip total | $24,995.46 | $24,995 |
| Gross profit | $2,789.25 · 12.4% | $2,789 · 12.4% |
| Every line after repeated edits | Unchanged | Correct to the cent |
| Per-section margins | 15.3% · 1.2% · 15.3% | Not shown on screen |
| What we tested | Lucas | Diego | Giancarlo |
|---|---|---|---|
| Every figure correct | Yes | Yes | One line overridden |
| Shows a section losing money | Yes | Calculated, not shown | Reports it as healthy |
| Over budget: warns and allows | Yes | Warns, then blocks | Yes |
| Service below cost: flagged | Yes | Yes | Prevented instead |
| Service with no rate: total marked | Yes | Yes | Shown as complete |
| Edits survive a reload | Wired up | Wired up | Not attempted |
| Tests and documentation | Yes | None | None |
| Matches your visual language | Clear | Closest | Solid |
His is the best written of the three. It names your actual clients, quantifies the overage and offers three specific fixes. It then stops the consultant from continuing. One of his three suggested fixes is to raise the agreed budget, and that is not possible anywhere in what he built.
Our view is that the software should warn loudly and let your consultant decide, because sometimes going over is deliberate. This is a design opinion rather than a mistake, and it is worth hearing his reasoning.
Two hotel rooms booked for the same nights, for two travellers. It was not in our answer key and nobody else noticed. He built a warning for it and used it to give the consultant a way to resolve the unpriced service.
Worth being precise, because this is his only real gap. His pricing engine calculates the per-section figures correctly, and the code even notes that a simple average of the markups "would lie." He computes them and never displays them. It is a presentation gap rather than a misunderstanding, and a short piece of work to fix.
Read line by line, and clean. The pricing formula lives in one place, and nothing in the interface recalculates money independently, so a row and a total cannot disagree. The auto-save handling is the most carefully built part of his submission, covering rapid edits, failed saves, retries and network interruptions, none of which was asked for. He shipped no automated tests.
The short version: a genuine alternative. If design quality and ease of communication matter more to you than $2,600 a month, he is the pick.
Senior front-end engineer at Bucksense since late 2023, and before that full-stack at Veckta, an early-stage startup where he worked directly with one of the founders from the beginning. Concrete results at both, and both on his CV: he reduced a page payload from around 19MB to 500KB, cutting worst-case load times from 30 seconds to 5, and built a caching layer that reduced search infrastructure costs by roughly 66%.
He builds the Abercrombie & Kent website, so itineraries, pricing, availability and how that information has to come together are familiar ground rather than something he would learn on your time. At Veckta he built an interactive modelling tool where changing an input immediately showed the financial impact, which is structurally the same problem as your pricing screen.
This is where he beats Lucas. His warning states are visually differentiated instead of uniform, his hierarchy is stronger, his contrast is better judged, and his screen sits closest to your existing product. He also shipped a mobile layout nobody asked for. For a founder with strong visual taste and no designer, that is not a small thing.
He ran his own electrical engineering practice for nine years before switching careers. It shows in how he explains technical decisions. He is the clearest of the three and the easiest to follow.
He can start now, and asks for a two-week trial before deciding whether to leave his current role. We would suggest taking him up on it. It is a bounded commitment that answers more than an interview does.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Peru · runs his own software startup · the most shipped client projects of the three
A capable engineer who has shipped more client work than either of the others. On this challenge he made one decision that this particular screen cannot carry, and missed two of the required pieces.
Available Immediately
The short version: his screen changes a price your consultant set, and the knock-on effect is that it reports a budget problem that does not exist.
The data includes one flight your consultant discounted by 15%, deliberately, to keep the trip under the client's budget. Giancarlo's system refuses that discount and quietly replaces it, charging the client $638 more. The row still displays the consultant's stated reason for the discount, immediately beside the number that ignores it.
| What we checked | Correct answer | His screen |
|---|---|---|
| Client price | $24,995 | $25,633 |
| Against the budget | $205 under | $433 over |
| Transport section margin | 1.2% | 12.2% |
| Total when a service is unpriced | Marked incomplete | Presented as complete |
The trip is not over budget. His override put it there, and his screen then advises cutting services to fix a problem it created. It also hides the section that is genuinely losing money.
| What we tested | Lucas | Diego | Giancarlo |
|---|---|---|---|
| Every figure correct | Yes | Yes | One line overridden |
| Shows a section losing money | Yes | Calculated, not shown | Reports it as healthy |
| Over budget: warns and allows | Yes | Warns, then blocks | Yes |
| Service below cost: flagged | Yes | Yes | Prevented instead |
| Service with no rate: total marked | Yes | Yes | Shown as complete |
| Edits survive a reload | Wired up | Wired up | Not attempted |
| Tests and documentation | Yes | None | None |
| Matches your visual language | Clear | Closest | Solid |
Nine of ten lines were priced correctly using the right definition of profit. He built the per-section margins that Diego's screen does not show. He made the correct call on the budget question, warning rather than blocking, where Diego did not. And his answer on repricing for more travellers was the best of the three.
See the code he wroteThe short version: a competent engineer whose instincts do not fit this particular screen. On a different product he would read very differently.
Giancarlo runs his own software startup and works across web and blockchain products. He has the highest count of confirmed shipped projects of anyone we screened in the technology your product uses, with named client work, repeat clients and strong public ratings.
You set pricing policy and your consultants make the commercial calls. An engineer whose instinct is to have the software overrule both is the wrong fit for the one screen where the company makes or loses its money. It is a correctable instinct rather than a limit on his ability, and correcting it costs a conversation.
What keeps him below the other two is not the instinct on its own. It is that alongside it he left out two required pieces of work. Three misses on the screen your business runs on is a lot to absorb, and the two above him did not have any.
Because the difference between his submission and the two above is the clearest illustration of what the challenge was designed to find. All three screens look competent at a glance. Two of them tell the truth about the money.
The dotted outline is what we think the role needs. These positions are our judgement, offered as a summary of the assessment above.
Both scored 9, so the number alone will not decide this. Here is what sits underneath it.
Lucas gives you accuracy, care and the lower rate
He is the most careful engineer of the three. His figures are exact, his screen is the only one that reveals a section quietly losing money, and he left behind tests and documentation that protect the work from whoever touches it next. He also trusts your consultants to make the commercial call. His design sense is good without being exceptional, and he is the hardest of the three to follow on a call.
Diego gives you design and smoother communication
He is the strongest designer here, with better hierarchy and contrast, closest to how your product already looks, plus a mobile layout nobody asked for. He is the clearest communicator of the three and already works in travel. He is equally accurate. He wrote no tests, left the section margins off the screen, and his version stops a consultant who wants to exceed a budget on purpose.
| What differs | Lucas · $30/hr | Diego · $45/hr |
|---|---|---|
| Quotes come out correct | Yes, every figure | Yes, every figure |
| UI design and visual polish | Clear and well organised | Noticeably more refined |
| Spoken English | Thick accent, takes effort | Easy to follow |
| General communication | Excellent in writing | Strongest of the three |
| Warns when part of a trip stops making money | Yes | Calculated, not shown |
| Lets a consultant go over budget deliberately | Yes, with a clear warning | No, stops them |
| Protects against future breakage | Tests and documentation | None written |
| Travel industry experience | No | Yes, builds A&K's site |
| Cost per month | $5,200 | $7,800 |
The dotted outline is what we think this role needs. These positions are our judgement, offered as a summary of the assessments above.
We would pick Lucas. Both of them build a screen that gets the numbers right, so the question is what the extra $2,600 a month buys. It buys better design and an easier conversation, which are real things. On the other side, Lucas is more thorough underneath, including the tests and documentation that keep a screen correct as it changes, and the saving could fund design help later if you want it. This is a close call rather than a formality. If meeting them shifts your view on design or on how easily you work together, Diego is a strong choice in his own right.
We have assessed the technical work and how each of them operates. What we cannot judge from the outside is whether you will enjoy working with this person, and on a two-person team that matters as much as anything in the scores.
Notice whether you enjoy the conversation. You will be speaking to this person most weeks for a year. Liking them is a real hiring criterion.
Notice how easily you follow them. Much of this work happens in writing, so it matters less than it would in an office. It is still the main thing separating our top two product engineers.
Notice whether they ask you anything. People who do well without a product manager tend to interrogate the problem rather than wait to be told what to build.
We covered the technical ground and how each person works. These are the questions we left for you, because the answers are about fit with you specifically: how you like to lead, how much structure you want to give, and who you want in the room when something goes wrong.
"Tell me about a time you disagreed with a founder or a client about how something should work. What happened?"
You have no technical co-founder. You need someone who will tell you when you are wrong, and do it well.
"Show me something you have built that you are proud of, and tell me why."
Reveals what they actually value: craft, correctness, speed or impact. You will know within a minute whether their taste matches yours.
"What does a week where you did your best work look like?"
Open enough that they cannot give you the answer they think you want. Tells you how much structure they need and how they prefer to communicate.
"What do you need from me to do your best work?"
You are not going to write detailed specifications. Better to find out now whether they expect them.
"When you do not know the answer to something, what do you do?"
With nobody technical beside you to catch a bluff, this matters more for you than for most founders. Listen for whether they are comfortable saying they do not know.
"What would make you want to leave in six months?"
Almost nobody asks this. The answers are usually honest and occasionally decisive.
Just say which of them you would like to meet and it will be arranged. Send over a calendar link and it can go straight to them, or email introductions work equally well, whichever is easier for you.
Happy to talk any of this through first, or to share the full technical assessment behind any candidate here.