返回 7*24 快讯
来源MarsBit

Can AI replace financial analysts? Vals AI's new version fails miserably in tests, with GPT 5.5 barely achieving over 50% accuracy.

According to Beating, an AI testing organization, Vals AI released its second-generation Financial Agent benchmark test (Finance Agent v2). This is an end-to-end test simulating the workflow of a junior financial analyst, containing 927 expert-reviewed questions. The new version of the test is significantly more difficult, with GPT 5.5 only achieving a 51.76% accuracy rate, extremely close to Claude Opus 4.7 (51.51%) and Claude Sonnet 4.6 (51.03%). Unlike single-round question-and-answer tests, this test requires the model to autonomously find relevant passages in hundreds of pages of 10-K and 10-Q financial statements, handle cross-year financial statement adjustments, and complete multi-step calculations with precise intermediate figures. Vals AI revealed that if a strict scoring standard of "must answer completely correctly" is adopted, the accuracy of all cutting-edge models drops below 40%; in the most difficult categories of "financial modeling" and "precedent analysis," the highest score is only 23%. In terms of other models, Kimi K2.6 ranked fifth with 44.87%, making it the highest-scoring domestic model; followed by GLM 5.1 (44.79%) and DeepSeek V4 (44.08%). Furthermore, the official "fastest speed" label was awarded to Claude Opus 4.7 (single time 360 seconds), while GLM 5.1 received the "most cost-effective" label (single cost $0.62). The collective drop in scores in this test (Opus 4.7 scored 64.4% in the previous generation test) proves one point: current AI can handle simple retrievals, but in the complex areas of finance where specific industry practices and extremely high numerical accuracy are required, it is far from replacing human analysts.
免责声明:以上内容仅为作者观点,不代表 711BTC 的任何立场,不构成与 711BTC 相关的任何投资建议。

相关推荐

9undefined前重要

153个被盗地址含132.95枚BTC,研究人员仍无法重现Coldcard攻击者种子

Odaily星球日报讯 据 Bitcoin News 监测,@PraveenPerera 发布的新研究显示,Coldcard 攻击者似乎先识别出存在漏洞的地址,再按照地址持有比特币数量排序,并从持仓量最高的地址开始分批转移。实际使用的转移软件则较为粗糙。一个地址拥有 225 个可花费 UTXO,攻击者恰好提取了最新的 200 个,留下最早的 25 个,其中包括一个价值 0.16 枚 BTC 的 UTXO。这与研究人员调查的一款区块链 API 默认返回 200 条记录的限制完全一致,表明攻击者可能未能加载下一页数据。该软件甚至花费了一个 294 聪的 UTXO,据报道使交易手续费增加约 2040 聪,花费金额明显高于该 UTXO 本身价值。研究作者据此认为,该工具的构建者对账户余额系统的理解可能强于对比特币 UTXO 模型的理解。尽管攻击者似乎已经获取受害者的完整种子,但至少 75 枚 BTC 仍留在由相同种子派生的其他地址中。目前最大的疑点是,153 个被盗地址中仍有 132.95 枚 BTC,研究人员仍无法复现这些地址背后的种子,因此不排除攻击者获取了未公开的私有设备数据或候选数据。

21undefined前

特朗普对多国无人机及零部件加征 15% 关税

ChainCatcher 消息,白宫宣布,美国总统特朗普将对来自欧盟、日本、列支敦士登、韩国和瑞士的输美无人机及零部件加征 15% 从价关税。

24undefined前

Reddit 将被纳入标普 500 指数

ChainCatcher 消息,Reddit (RDDT.N) 将被纳入标普 500 指数。

25undefined前

巴尔的摩市起诉Kalshi与Polymarket,指控经营无牌体育博彩平台

Odaily星球日报讯 巴尔的摩市与市长 Brendan Scott 已就 Kalshi 和 Polymarket 提起诉讼,指控两家公司违反当地赌博法律及欺骗性商业行为规定。市长办公室周四表示,两家公司运营“非法、无牌体育博彩平台”,并误导用户了解产品的合法性及监管状态。 巴尔的摩市方面认为,Kalshi 和 Polymarket 将事件合约描述为交易,但相关交易实质上属于州法律禁止的非法博彩。针对 Kalshi 的诉状还将 Robinhood、Webull 和 Coinbase 列为其预测市场平台合作伙伴,并指控相关公司将体育合约宣传为可在马里兰州合法购买和交易的产品。 美国商品期货交易委员会(CFTC)及相关公司认为,预测市场事件合约属于其监管范围内的“互换”。Polymarket 表示,在 CFTC 注册交易所运营的预测市场受联邦法律管辖,不应适用州及地方层面的规则。(Cointelegraph)

35undefined前

总名义价值达2.02亿美元,Bowen_bug在markets.xyz下达40笔SPCX空单

Odaily星球日报讯 据 HyperliquidNews 监测,markets.xyz 上名为 Bowen_bug 的钱包下达 40 笔 SPCX 空单,每笔涉及 3.55 万枚 SPCX,单笔名义价值约 505 万美元,总名义价值达 2.02 亿美元,订单价格位于 142.43 美元至 142.9 美元之间,价差为 0.33%。相关信息由@mlmabc 提供。

08-14 04:36

惠誉确认美国信用评级为 “AA+”

ChainCatcher 消息,惠誉确认美国信用评级为 “AA+”,展望稳定。