LIVE AT ARCO.CAREERS · BUILT SOLO WITH AI

One designer. No engineering team. A live product.

Arco is an AI job-search platform — a scored feed, a coach, and a pipeline, live at arco.careers. It's also a working demonstration of how I design now: every feature specified as a structured prompt, executed through Claude Code, and verified by me on the live deploy.

This case study documents the method as much as the product — including where it broke.

Role: Product, design, engineering direction — solo · May 2026 → live · ongoing

Stack: React + FastAPI + PostgreSQL, shipped through Claude Code

May 2026 → live · ongoing

arco pipeline board with “New” and “Applied” columns of job cards, each showing the company, role, salary range and a fit score.

Honest scores over flattering ones

A fit score only works if users can trust a low one.

Coach, don't generate

Tailoring that starts from your real experience, not a template.

The pipeline is the product

From first match to offer in one place — feed, prep, and tracking connected.

Built with the tools it runs on

Designed and shipped solo with Claude Code and Figma MCP.

THE PRODUCT

The evidence, briefly

Job boards give you volume and no comprehension. Arco scores every role against your actual profile, explains the score instead of asserting it, and takes you from "who is this company?" to a tailored application — feed, coach, and pipeline in one place. I built it during my own senior-design-role search, which means every scoring decision got tested on the most demanding user available: me, with something at stake.

Role briefing for a Product Designer role at Figma with a fit score of 94: panels for the brief, verdict, compensation and a three-year compensation trajectory.
Every role opens as a briefing, not a listing: what the company actually does, what it pays, and why the score is what it is.

HOW EVERY FEATURE SHIPS

The loop

Every feature crosses the same cycle: a design decision, written as a scoped prompt; Claude Code executes it; I verify it on the live deploy. The loop ends at verification, not shipping — most AI-workflow diagrams end at "deploy," and that's the flattering version. What actually travels back around the arrow is regressions, drift, and misses.

A prompt in this system is a design artifact. It opens with reconnaissance — read the code before touching it. It scopes the change surgically. It names what must not be touched. And it closes with a deploy sequence that ends in me checking the live build hash, because "done" is a claim, not a fact.

The build loop: design decision, scoped prompt, Claude Code, verified deploy. Figma MCP carries design tokens across the seam, and regressions, drift and misses loop back so that every correction re-enters as the next prompt.
May 2026
Prompt from May 2026: “Can we make some cool animation on this get-started button? Maybe we can have the button in black and then have a green golden shimmer go from left to right?”
August 2026
Prompt from August 2026: “Step 1: audit first — don’t write code yet … Report back before changing anything.” — followed by an explicit do-not-touch list and: “Push, paste the Railway HEAD hash, then tell me how to trigger a manual crawl for my account so I can confirm.”

Three months apart. Every line of the second prompt was paid for by a specific incident.

THE JUDGMENT BOUNDARY

What the AI couldn't decide

The honest question about AI-built products is: what did the human actually do? The answer is everywhere the model needed a decision it had no basis to make — and everywhere it was wrong and didn't know it.

"If the user wants to start clicking around the screen, at least the banner is still there."

Claude specced a full-screen waiting takeover for the first crawl. I killed it. A takeover hijacks the app for minutes and strips the user of agency. The plan was rewritten around enhancing the banner in place — a call about user agency, not code.

"Not sure 72 is good enough? How come?"

The optimizer took my resume-match score from 62 to 72 and I pushed back. The system's answer held: the missing points were substance gaps it refused to pretend away. A tool that returned 95 would be lying, and lying resumes get found out in interviews. I kept the 72. That exchange became the product's core principle: the score's job is to be right, not to be liked.

"Fix this, it's the 4th time now lol"

A modal button kept coming back broken after each "fix." On the fourth attempt the model finally admitted it had been adjusting the wrong attribute the whole time. The loop breaks on human persistence, not model confidence. That's what supervision of AI work actually looks like.

None of these live in the code. They live in the taste, the pushback, and the refusal to accept a confident wrong answer — which is the job.

WHERE IT BROKE

Paid for in incidents

Every rule in my prompt structure exists because its absence cost me something. The method wasn't designed; it accreted.

The dist trap. The deploy pipeline serves a pre-built frontend from git. For weeks, the AI would edit source, commit, push — and the live product wouldn't change, because the build step was silently skipped. It took multiple "deployed but unchanged" incidents before the fix stopped being a reminder and became infrastructure: a pre-commit hook that mechanically blocks the failure, because rules the AI has to remember are rules it will eventually forget.

THE RULE — EVERY PROMPT ENDS WITH BUILD, STAGE DIST, PUSH, VERIFY THE LIVE HASH.

The migration that crashed three times. One database migration crashed production three times in a row — type mismatches the local test environment couldn't reproduce, then orphaned rows from deleted users. Each crash rolled back cleanly; the cost was downtime, not corruption. The doctrine that came out of it: read the real schema before writing a migration, assume dirty legacy data, and dry-run every migration against production inside a rollback before pushing. An engineer might have known this on day one. I bought it with three crashes.

THE RULE — READ THE REAL SCHEMA FIRST. DRY-RUN EVERY MIGRATION AGAINST PRODUCTION INSIDE A ROLLBACK.

The failures nothing catches. The worst class of failure throws no error. A one-word bug silently reported zero new jobs on every digest run. A hard-coded timeout was quietly cutting off the majority of scraping runs. Both shipped, both ran, and nothing in the loop caught them — until a number looked wrong to me ("2 jobs in total for global remote blockchain designer jobs is a bit low?") and a read-only audit traced it down. Verification is strong on what you can see. On what you can't, the last line of defense is a human who knows what the numbers should feel like.

THE RULE — WHEN A NUMBER FEELS WRONG, RUN A READ-ONLY AUDIT BEFORE ASSUMING IT'S FINE.

These rules now live in the repo itself — a CLAUDE.md the AI reads before working, and hooks that enforce what memory can't. The method's output isn't just the product; it's the guardrails.

AI-NATIVE BY DESIGN

Designed with AI, for AI

Every score costs money, and where the spending stops is a design decision. A full crawl pulls 250–350 raw roles; a location gate cuts that to roughly 115. Scoring all of them for every free user was unsustainable — so free crawls run a cheap heuristic pre-rank first and send only the top 50 candidates to the model, while paying users get every candidate scored. Everyday tasks route to a lighter model; judgment-heavy work — merging four CVs and resolving their contradictions — gets the stronger one. Where those seams sit is revisited constantly. The paywall stays honest by the same logic: the locked matches are real ranked results, not filler.

The same job once showed a 92 on the board and a 62 in the optimizer. Both were correct: one measured job fit, the other how well my resume made the case. The fix wasn't reconciling the numbers — it was naming them. "Job fit" and "Resume match" are now two labeled scores, and the gap between them is the pitch: a great-fit job your resume undersells is exactly what the coach exists for. Honesty is enforced in small places too. The first-crawl banner used to cycle fake timer-driven progress phases; it now reports real pipeline stages, and its time estimate was corrected against observed run data — "usually 5–10 minutes," because that's what's true. A "NEW" badge that mislabeled old jobs was a lie in miniature, and got the same treatment.

That includes designing for AI, not just with it: model selection, cost seams, prompt architecture, and the UX of machine judgment (how a score earns trust, how a coach corrects without condescending) are product design decisions here, as much as any screen.

A score that flatters everyone is a score nobody checks twice. Fewer, harder matches are what make the number worth reading.

Two flows: “everything’s a match” leads to “trust collapses”; “fewer matches” leads to “trust compounds”.

THE COACH

A different kind of space

The coach began as one tab inside a dense, multi-panel product — reading as "more app," not a conversation. Three treatments were mocked up: a quiet inset canvas (separation without identity), a slate-blue tint on the light surface, and a fully dark immersive coach that first got filed as "a big bet — maybe later." Two observations reframed it: the board's rail was already dark, making dark the consistency fix rather than the risk; and the treatment had to survive the panel expanding into a full-width workspace, which meant both ends of that animation needed the same surface.

The typography carries the same intent: the app runs on Geist, but the AI's messages are set in Georgia — a serif voice for the machine, so you always know who's talking. Tailoring is tracked per job — 78 → 85 — without ever touching your master CV: the machine's judgment is scoped, and the microcopy promises it.

arco AI Coach on a dark surface: a list of roles on the left and a conversation on the right with sections “Equity specifics” and “Quick wins to check”, set in a serif typeface.
The coach runs on its own dark surface, scoped to read as a different kind of space.
Each conversation is bound to one role, and tailoring is tracked per job — 78 → 85 — without ever touching the master CV.

OUTSIDE FORCES

The system doesn't control its world

Job sourcing is fragile by nature. Boards block crawlers, change structure without notice, or re-render as apps with nothing to parse. Two blockchain sources were decommissioned in a single June week — one behind a new bot wall, one rebuilt as a JavaScript app with no static feed — and the survivor is read through its sitemap instead of its pages, because that's the surface that doesn't break. A clean, deduplicated, correctly classified feed is continuous maintenance, not a one-time integration.

The strangest design constraint came from payments. A merchant-of-record's category review read Arco as a "resume builder" and a "job board" — two categories it wouldn't support — and declined checkouts. Both readings were wrong, but they were plausible from the outside, and that was the design problem: the product's surface was underselling its actual center. The fix was positioning, executed like any other feature: landing copy rewritten to lead with the AI matching and coaching layer, CV optimization demoted to a supporting beat, and a one-line counter-framing that now defines the product externally — Arco doesn't sell job listings, host a public board, or sell CV templates; the subscription is the intelligence layer. A payment processor forced the sharpest positioning work of the project.

LIMITS

What this method can't do

The honest version: this way of working has real edges. Failures that throw no error persist until a human notices output looks wrong. A confident model can iterate on the wrong attribute four times in a row. Production taught me database lessons an engineer would have known on day one. And the first surface you build is rarely the one you keep — Arco started mobile-first under another name; the desktop app overtook it, and mobile is now being rebuilt to match the product it became. That's the honest cost of learning in public.

No usage metrics on this page. Nothing is measured rigorously enough yet to put a number on, and I'd rather show real work than dress it up.

arco desktop app: the job board greeting “Hey, Sabrina”, a role briefing in the middle and the AI coach panel on the right.
Mobile came first, under an earlier name. Desktop overtook it — and the phone now has to be rebuilt to match the product it became.
arco mobile app: counts of 1051 active roles and 115 apply-now roles above job cards for Figma, FINQ Israel and Solana Foundation with fit scores of 94, 92 and 92.

WHERE THIS STANDS NOW

Built, open, next

Working today:

Arco is live at arco.careers. The scored feed, the coach — CV analysis, per-job tailoring, guided optimization — and the pipeline board run end-to-end, with free and Pro tiers defined and a public help center.

Still open:

Payments are in final activation with a merchant of record, so subscriptions aren't switched on yet. The mobile app is being rebuilt to catch up with desktop. Job sources are expanding lane by lane.

Next:

Payments live, mobile redone, sourcing broadened — in that order.

View Product

I design systems, not just screens—products that stay clear under real-world use.

Back to Top