iFANN
    iFANNを検索...
    ログイン
    ホーム
    ニュース
    動画
    写真
    GIF
    見つける
    投票
    アワード
    iFAMOUS
    ウィキ
    アニメ
    ルーム
    通知
    メッセージ
    ブックマーク
    プロフィール
    ウィキアワードiFAMOUSランキング業界クリエイター報酬ユーザー報酬利用規約プライバシーコミュニティガイドライン削除申請 / DMCAヘルプ開発者

    © 2026 iFANN

    ホーム
    検索
    メッセージ
    お知らせ
    プロフィール
    写真
    Nate
    Nate@nate_5122w
    💭Tech💭AI
    CommerceAgentBench Leaderboard

    @nate_512There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    元の投稿を見る

    CommerceAgentBench Leaderboard

    @nate_512さんの写真· Aug 31, 2026· Tech

    この写真について

    This is a screenshot of a leaderboard for AI models, specifically showing their pass rates on various real-world workflows. The focus is on the performance data presented in bar charts and tables. The mood is informative and analytical, with a clean, data-driven aesthetic. A notable detail is the ranking of different AI models like Claude Opus, GPT, and Gemini, with their respective pass rates displayed. The title at the top reads "CommerceAgentBench Leaderboard" and the Accio logo is visible in the top right corner.

    Techの写真をすべて見る

    ?

    Techの写真をもっと見る

    Techの写真をすべて見る
    PC graphics settings guidePC graphics settings guideKyle Sonlin live on tokenizationKyle Sonlin live on tokenizationWaymo car police arrestWaymo car police arrestLuke Rudkowski questions Jeff Bezos at jury dutyLuke Rudkowski questions Jeff Bezos at jury dutyData center e-waste projection2Data center e-waste projectionHermes local models one-click setupHermes local models one-click setup20 GrokBot Tips from poteto Founder Session20 GrokBot Tips from poteto Founder Sessionmuseum exhibit cartoonmuseum exhibit cartoonNanchang drone show at Qiushui Square4Nanchang drone show at Qiushui SquareAmelia SINDAZA CYIBARUMA FLORIEN reactionAmelia SINDAZA CYIBARUMA FLORIEN reactionLenovo Yoga 7i specs and price4Lenovo Yoga 7i specs and priceLenovo ThinkBook 13s sale4Lenovo ThinkBook 13s sale2026 Atlantic hurricane season forecast2026 Atlantic hurricane season forecastAI generates photo of a womanAI generates photo of a womanWalmart Tap to Pay CaliforniaWalmart Tap to Pay CaliforniaBest Electronics Brands by CategoryBest Electronics Brands by CategoryVanessa financial constraintVanessa financial constraintPOLSIA founder story strategyPOLSIA founder story strategy
    写真
    Nate
    Nate@nate_5122w
    💭Tech💭AI
    CommerceAgentBench Leaderboard

    @nate_512There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    元の投稿を見る

    CommerceAgentBench Leaderboard

    @nate_512さんの写真· Aug 31, 2026· Tech

    この写真について

    This is a screenshot of a leaderboard for AI models, specifically showing their pass rates on various real-world workflows. The focus is on the performance data presented in bar charts and tables. The mood is informative and analytical, with a clean, data-driven aesthetic. A notable detail is the ranking of different AI models like Claude Opus, GPT, and Gemini, with their respective pass rates displayed. The title at the top reads "CommerceAgentBench Leaderboard" and the Accio logo is visible in the top right corner.

    Techの写真をすべて見る

    ?

    Techの写真をもっと見る

    Techの写真をすべて見る
    PC graphics settings guidePC graphics settings guideKyle Sonlin live on tokenizationKyle Sonlin live on tokenizationWaymo car police arrestWaymo car police arrestLuke Rudkowski questions Jeff Bezos at jury dutyLuke Rudkowski questions Jeff Bezos at jury dutyData center e-waste projection2Data center e-waste projectionHermes local models one-click setupHermes local models one-click setup20 GrokBot Tips from poteto Founder Session20 GrokBot Tips from poteto Founder Sessionmuseum exhibit cartoonmuseum exhibit cartoonNanchang drone show at Qiushui Square4Nanchang drone show at Qiushui SquareAmelia SINDAZA CYIBARUMA FLORIEN reactionAmelia SINDAZA CYIBARUMA FLORIEN reactionLenovo Yoga 7i specs and price4Lenovo Yoga 7i specs and priceLenovo ThinkBook 13s sale4Lenovo ThinkBook 13s sale2026 Atlantic hurricane season forecast2026 Atlantic hurricane season forecastAI generates photo of a womanAI generates photo of a womanWalmart Tap to Pay CaliforniaWalmart Tap to Pay CaliforniaBest Electronics Brands by CategoryBest Electronics Brands by CategoryVanessa financial constraintVanessa financial constraintPOLSIA founder story strategyPOLSIA founder story strategy