I benchmarked Claude Sonnet 4.6, Gemini 2.5 Pro, and Qwen3 Max Thinking on translating bilingual kitchen transcripts into valid culinary JSON schemas. Here is how they handle complex schemas, Konglish loanwords, and cultural nuances.
I benchmarked OpenAI gpt-4o-transcribe and Deepgram Nova-3 on real Korean/English kitchen audio — clean and noisy. Here are the CER, WER, latency, and cost numbers so you can pick the right model without guessing.