Spreadsheets are the obvious choice when you're starting to evaluate data. They're easy to set up and work fine for tracking labels, annotators, and progress. But spreadsheets eventually reach their ceiling when you scale: every team keeps its own conventions, QA turns manual and inconsistent, and there's no version control or shared workflow to tie it all together. Past a certain point, that setup can't keep up with the volume and variety of LLM evaluation work.
GetYourGuide hit that wall as Gen AI took on a bigger role across the business. By replacing a patchwork of team-owned spreadsheets with a centralized annotation platform, GetYourGuide onboarded 45 internal users across data science, product, and engineering in a single month and processed more than 340,000 evaluation items in their first year.
Evaluation stopped being the bottleneck, and the team could ship new generative features to travelers much faster.
The Challenge
At GetYourGuide, evaluation is tied directly to product quality. It's what lets them ship generated recommendations, ranking algorithms, and customer-facing chatbots that customers can rely on. Each model output is measured against quality, safety, and alignment criteria before it ships, and reviewed continuously after it does.
As Gen AI showed up in more parts of the product, evaluation had to keep up so the team could move fast without losing consistency or accuracy.
But the way they handled evaluation wasn't built for the volume and variety of LLM use cases now coming through.
- Each team kept their own spreadsheets, with bespoke columns, ad-hoc conventions, and minimal automation.
- Tasks ranged across multi-turn chat evaluation, text classification, entity extraction, and free-form LLM response grading — but no shared workflow connected them, so a lot of time went into manual coordination.
- Version control was unreliable.
- Traceability between an evaluation result and the model version or prompt that produced it was fragile.
- Quality assurance was inconsistent from team to team.

With LLM use set to keep growing across the company, spreadsheets weren't going to hold up over the next 12–24 months. GetYourGuide needed an evaluation setup built for LLMs, flexible enough to grow with new use cases, and easy to use for both technical and non-technical people.
Platform Criteria
Before selecting a partner, the team identified criteria to support their growing needs. Ting Wang, Director of Engineering, knew the platform had to:
- Handle the data types and evaluation patterns they were already working with – chat, text, classification, structured outputs, LLM response grading.
- Let non-engineers build and edit custom workflows in the UI, without engineering involvement.
- Bring technical reviewers and product/ops people into the same tool.
- Cover the whole project lifecycle – managing annotators, automating workflows, running QA, versioning, and tracking experiments.

And it had to fit where GetYourGuide's LLM strategy was going, not just where it was at the time.
Why SuperAnnotate
SuperAnnotate stood out as the clearest fit on every dimension. In Ting Wang’s words,
“SuperAnnotate has been a fantastic upgrade for our data labeling efforts. Moving away from labeling datasets in spreadsheets to a purpose-built tool reduced errors and made the whole process far more reliable. Their project tracking and analytics dashboard gave us valuable data points not only to assess projects on completion, but also to make important adjustments along the way and ensure each one landed on time. We also appreciated the intuitive UI builder in the project creation phase, which was accessible enough that even non-technical staff could set up projects exactly the way we wanted. Overall, we're really happy with the efficiency and quality SuperAnnotate delivered, supporting our ongoing and growing data labeling needs at a crucial time for our teams.”
- Ting Wang, Director of Engineering, GetYourGuide

Ultimately, SuperAnnotate’s operational layer allowed GetYourGuide to unlock annotator management, QA workflows, versioning, and traceability. This had not previously been possible in spreadsheets.
Solution & Results
With SuperAnnotate, GetYourGuide moved from scattered, spreadsheet-based evaluation to a single platform the whole team could work in.
Implementation focused on three things:
- Standardizing evaluation schemas across use cases,
- Consolidating teams onto common workflows,
- Replacing manual handoffs with platform-managed processes.

Led by SuperAnnotate’s best practices, the GetYourGuide team built custom evaluation schemas for each major LLM use case, then set up multi-step review and QA pipelines. Rather than manual handoffs, annotators received the right items at the right stage with the right context, automatically. Version control and dataset tracking are built in, so the team can trace any evaluation result back to the exact model, prompt, or dataset that produced it.

One of the benefits was the cross-team collaboration that was enabled alongside the tooling. Data science, product, and engineering colleagues began working from shared datasets and pipelines instead of one-off files, with visibility into who was reviewing what, where bottlenecks were forming, and how quality compared across batches.
Most importantly, evaluation has stopped being a bottleneck and started being a multiplier. New LLM use cases can plug into shared workflows instead of inventing their own. New annotators can be fully onboarded within days.
Evaluation is now a core part of how GetYourGuide builds with AI, not just a step in the process. The team can roll out new LLM features knowing the quality checks behind them hold up, and keep adding new GenAI use cases over the next 12–24 months without rebuilding the evaluation setup every time.
About GetYourGuide
GetYourGuide is a global travel experience platform that connects travelers with tours, activities, and unique experiences around the world. The company has made significant investments in product innovation, and is at the forefront of applying AI and data across key pillars of the business, including traveler experience, supplier tooling, and marketing. GetYourGuide’s Data Product and machine learning engineering teams support critical efforts across the organization including customer-facing AI models that are deployed in production to enhance discovery and personalization. They also develop tools to simplify supplier onboarding and improve data quality such as LLMs that convert suppliers’ free-form text into structured data and have a dedicated AI team focused on marketing areas such as paid search optimization, and CRM.
About SuperAnnotate
SuperAnnotate is the enterprise AI data platform that helps the world's leading AI teams build efficient human data and evaluation pipelines to ship better agentic, multimodal, and frontier AI. The platform combines flexible annotation and evaluation tooling with human and agent-in-the-loop workflows, advanced quality control, and the option to tap into SuperAnnotate's expert talent network when teams need to scale beyond their internal capacity. SuperAnnotate is trusted by frontier AI labs and applied AI teams around the world, and is backed by NVIDIA.
Want to see how SuperAnnotate could work for your team? Book a demo.


