Ask Lonic

What would you like to know?

Answers are drawn from Lonic's published reporting on lonic.bond, with every source listed.

No account needed — answers are generated from our article library.

Answer

computer use

Several frontier labs have claimed state-of-the-art results on 'agentic computer use' benchmarks in recent weeks, each measuring how well a model can operate a computer the way a person would: opening applications, navigating menus, filling in forms, moving files between programs, and completing multi-step tasks across a graphical interface rather than through a dedicated API. The claims are real in the narrow sense that the reported scores are usually reproducible on the specific benchmark suite in question. What they tell you about deployment-ready reliability inside an actual company is a considerably smaller thing than the leaderboard position implies.

  • Tasks are drawn from a known, finite pool, which creates an incentive to fine-tune specifically on patterns resembling the benchmark rather than on general interface competence.
  • Sandboxed environments are simplified and stable compared with production software, which regularly changes layout, introduces pop-ups, or behaves inconsistently in ways the benchmark never models.
  • Automatic scoring against an expected end state can reward an agent that reaches the right screen through a fragile or unsafe sequence of actions that happened to work once.
  • A single aggregate score obscures wide variance in performance across task categories, so a strong headline number can coexist with poor performance on the categories most relevant to a given deployment.

People also asked

Browse the whole library

New here? Start with today's trending stories or read how Lonic reports.