---
type: idea
title:
created: 2026-07-07
status: seed
origin_sources: []
related_auto: []
used_in: []
---

## Claim
- Okay, as somebody working in strategy and finance and working with AI products or model the products that work on top of frontier models and with the recent conversation or rhetoric around frontier intelligence costs and the rise of open source, I've been diving deep on understanding or, kind of just educating myself on how to assess the differences between different models and be able to judge or at least hold space or hold, some influence in a room where conversations are happening around model choice. I have seen technical folks be extremely bullish on using frontier intelligence, but it takes a business voice to push back sometimes and find a middle ground. That makes me want to learn more about creating and running evals and actually reading the outputs. Over the last week, I created a GLM 5.2 worker and used my Zippo AI GLM coding plan to run a set of evals designed to measure ambiguity, false positivity, etc. on finance workflows because those are workflows I understand best and how codecs and GLM 5.2 varied in their approaches.
- I ran a robust eval test on a bunch of different tests comparing codex and GLM 5.2. Of course I worry a lot of the cost of intelligence, somehting I think about and work on at Epic Games extensively. 
- One thing that I find frontier models were actually better at is the real world usability
- We dont prompt perfectly, we don't offer perfect context, sometimes we just expect the models to know what we want - what would be obvious to a  
-


- Recently, there has been a rise in the rhetoric around increasing token costs and the ROI narrative on frontier intelligence has gain more and more ground. My job requires me to weigh in on these conversations. 
- To be very honest, until very recently, I did not take open-weights models as real alternatives to frontier models. GLM 5.2 along with Fable release-and-pull-back, Satya's post on owning enterprise context, push by inference providers to innovate on OSS inference offering lightning speeds and costs - all came together at a time that really challenged frontier labs' revenue durability. 
- 