Eu li um comentário no HN que me chamou atenção (negativamente):
(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science
O segundo parágrafo, falando sobre a escrita, pegou mal para mim. Um modelo tão caro chamar atenção porque "escreve bem" faz parecer que não chamou atenção em coisas mais interessantes. E me lembro de o Claude começar a se popularizar no início do ano porque "escrevia bem", mas eu não acho que a escrita piorou, apenas acho que todos se acostumaram, porque todos estão usando.
Da mesma forma que é fácil abrir uma publicação e ver que foi escrita por IA, se esse novo estilo se popularizar em outros modelos, ou muitas pessoas usarem o Fable para isso, logo parecerá ruim novamente.