The website audit that runs on www.ranbanka.com is one of our own AI systems. A prospect submits a URL and receives a branded email report covering performance, mobile, UX and conversion, SEO and code quality. We built it for ourselves, we run it in production, and every audit it produces costs us money.
The first version was a CrewAI crew. It worked, but it was slow, it invented scores, and it broke in ways that were hard to predict. This case study explains our CrewAI to LangGraph migration: what was wrong with the original design, what we replaced it with, and what changed once the new pipeline was live. We treat it as proof of the kind of AI system we can build for clients. Our AI solutions page describes how we approach that work.
The challenge
The CrewAI version had five problems, and each one affected speed, honesty or cost.
- Repeated work. Six agents ran in sequence, and each one fetched the page and the PageSpeed data again. The same network calls happened over and over in a single audit.
- Guessed scores. The LLM produced every score itself. Nothing tied the numbers to an actual measurement.
- Fragile output. The report was free-form Markdown, parsed with regular expressions. When the model's formatting drifted, the parsing failed and the dashboard in the email simply disappeared.
- Made-up reports. If a site could not be reached, the crew still produced a report. That report was fiction.
- Slow runs. Each audit took several minutes.
Because the audit costs us money on every run, we also needed hard limits on how often it could run, without letting simultaneous form submissions slip past those limits.
What we built
We rebuilt the audit as a LangGraph pipeline with three clear stages: gather, analyse and assemble. Each stage has one job.
Gather: one fetch, no LLM
The pipeline fetches the page's HTML once and calls the Google PageSpeed Insights API for mobile and desktop in parallel. An HTML signal parser then extracts the evidence the analysts need: title, meta tags, headings, image alt text, scripts, calls to action and headers. This stage uses no LLM at all.
If the page cannot be fetched, the run stops here, before any LLM call is made. The visitor receives a polite "we couldn't reach your website" email instead of an invented report.
Analyse: five analysts in parallel
Five analysts run at the same time: Performance, Mobile, UX/Conversion, SEO and Code Quality. Each is a single structured-output Claude call that receives only the evidence its category needs, not the whole page. Structured output means each analyst returns data in a defined shape rather than free text, so the report no longer depends on parsing Markdown with regular expressions. Each call gets one retry on failure.
Assemble: real scores, fixed order
The assemble step uses real Lighthouse scores from PageSpeed Insights for Performance and SEO, and the analysts' scores for the other three categories. It builds the report in a fixed order and turns it into a branded HTML email for the prospect, sent through Amazon SES. With a fixed structure, the email dashboard always renders.
How we delivered it
The pipeline was built in-house by the Ranbanka team on this stack:
- LangGraph and Python for the pipeline itself, with parallel branches for the PageSpeed calls and the five analysts.
- Claude Sonnet 5.5 for the five structured-output analyst calls.
- AWS Lambda, packaged as a Docker image on Amazon ECR. The contact-form Lambda starts the audit asynchronously.
- Google PageSpeed Insights API for real Lighthouse data on mobile and desktop.
- Amazon SES for the branded email report.
- DynamoDB for the usage limits described below.
The hard part: fast, honest and affordable
The biggest challenge was making the audit fast, honest and affordable enough for us to keep offering it without charge. Speed came from fetching once and running the analysts in parallel. Honesty came from using real Lighthouse scores where they exist, giving each analyst only relevant evidence, and refusing to report on sites that could not be reached.
Cost control needed more care. The system allows at most five automatic audits per day, and it also checks for repeat sites and repeat email addresses. These checks are enforced atomically in a DynamoDB transaction, so concurrent submissions cannot push the count past the limit, even if several arrive at the same moment.
Safe rollout
The pipeline is covered by an offline test suite. We also kept the previous CrewAI image, so a rollback option stays available if it is ever needed.
The result
Measured on www.ranbanka.com:
- About 45 seconds per audit, down from several minutes.
- About $0.085 per audit (shown on our portfolio as ~$0.09), from five Claude calls.
- Real Lighthouse scores for Performance and SEO, instead of numbers guessed by a model.
- An email dashboard that always renders, because the report is built from structured data in a fixed order.
- No invented reports for unreachable sites: the run stops before any LLM call.
- A daily cost cap with repeat-site and repeat-email checks enforced atomically, so concurrent submissions cannot overshoot it.
The broader lesson for anyone weighing a CrewAI to LangGraph migration: much of the improvement came from taking work away from the LLM. Fetching, parsing, taking scores from Lighthouse and assembling the report are all plain code. The model is used only where judgement is needed, with narrow inputs and a defined output shape.
You can see this alongside our other AI and frontend projects in our portfolio, and the full range of what we offer on our services page.
Planning an AI pipeline of your own?
If you have an agent workflow that is slow, costly or unreliable, or you are planning a new AI system and want it built with clear cost limits and verifiable output, talk to us. We respond within 24 hours, work under NDA, and the initial consultation is free.
