<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>manapick AI 更新フィード</title>
  <link>https://ai.manapick.app/</link>
  <atom:link href="https://ai.manapick.app/feed.xml" rel="self" type="application/rss+xml" />
  <description>manapick AIのニュース・ガイド更新をまとめたRSSフィードです。</description>
  <language>ja</language>
  <lastBuildDate>Fri, 07 Aug 2026 15:00:00 GMT</lastBuildDate>
  <ttl>60</ttl>
  <item>
    <title>GitHub Copilotコードレビュー、軽量・標準の深さを正式提供——PRごとと組織全体で選択</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-github-copilot-review-effort-ga/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-github-copilot-review-effort-ga/</guid>
    <description>定型変更はLite、複雑・高リスクな変更はBalancedを選択可能。組織の既定値と実行履歴の表示にも対応した。 小さなPRと重要なPRでAIレビューの深さを分けられますが、テストと人の確認は引き続き必要です。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>GitHub、Claude・Codexなどエージェント別の利用数をAPIで集計——1日・28日レポートに追加</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-github-agent-app-metrics/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-github-agent-app-metrics/</guid>
    <description>Copilot利用指標APIが外部エージェント別の起動数とセッション数に対応。導入後の利用状況を比較しやすくした。 会社で導入したAIエージェントの利用実績を比較できますが、集計対象外と項目定義を確認して読む必要があります。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>work</category>
  </item>
  <item>
    <title>GitHub Copilot、企業向けMCP許可・拒否リストを正式提供——不明な設定は安全側で遮断</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-github-copilot-mcp-allowlists/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-github-copilot-mcp-allowlists/</guid>
    <description>URLやローカルコマンドでMCPサーバーを一括管理。Copilot app、CLI、VS Codeに適用し、設定不備は接続を止める。 会社で使うMCP接続先を一括制御できますが、許可後も権限・ログ・サーバーの安全性を継続確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>Gemini Notebooks、資料の自動追加に対応——Drive・YouTube・Webを定期ワークフローで更新</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-gemini-notebooks-auto-sources/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-gemini-notebooks-auto-sources/</guid>
    <description>Workspace Studioに資料追加ステップを新設。テキストや各種リンクを定期的に取り込み、ノートの更新を自動化する。 定例資料をノートへ自動で集められますが、追加された情報の日付・権限・正しさは人が確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Gemini Notebooks</category>
    <category>work</category>
  </item>
  <item>
    <title>arXiv論文：AIスキルの再利用を3種類の痕跡で監査——3万6446件を調査、F1は0.898</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-skilltrace-provenance/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-skilltrace-provenance/</guid>
    <description>文章・実装・作業手順を別々に照合するSkillTraceを提案。コードだけでは見逃す再利用の根拠を示す。 AIスキルの出典確認を広げる方法ですが、類似結果だけで盗用と決めず、履歴とライセンスを確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>研究論文</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：調査AIの長い失敗記録を自動監査——1243件、平均6万5100トークンで原因を特定</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-searchauditor/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-searchauditor/</guid>
    <description>検索AIの誤り箇所・原因・直し方を評価するSearchAuditBenchを公開。提案法の総合合格率は32.3%。 調査AIの途中の間違いを見つける研究ですが、提案法も合格率32.3%で、人の出典確認はまだ欠かせません。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>研究論文</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：音・映像・文章を扱えるAIも矛盾に弱い——人88.64%、最良モデル73.17%</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-c3po-omnimodal/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-c3po-omnimodal/</guid>
    <description>4種類の情報を組み合わせるC3PO評価を提案。失敗の86〜95%で一つの形式に偏る現象を確認した。 複数形式に対応したAIでも文章だけに偏る場合があります。重要な判断では音声・映像・画像の根拠も個別に確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>研究論文</category>
    <category>search</category>
  </item>
  <item>
    <title>GitHub Copilot更新：CLIに/worktreeと/rewind——並行作業と巻き戻し、VS Codeは横質問に対応</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-github-copilot-weekly-aug3/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-github-copilot-weekly-aug3/</guid>
    <description>Copilotアプリ、CLI、VS Code 1.132の週間更新。作業を止めずに分岐・復元・質問しやすくなった。 AI開発を分岐・巻き戻しできるため試行錯誤しやすくなりますが、実験機能の差分は必ず確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>GitHub Copilot、費用とPR数を並べるROI欄を追加——給与帯を変えて試算、因果関係には注意</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-github-copilot-roi-dashboard/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-github-copilot-roi-dashboard/</guid>
    <description>Impact dashboardに「潜在的な投資対効果」を追加。AIクレジット費用、給与比率、PR数を利用段階別に比較する。 Copilotの費用と開発量を同じ画面で試算できますが、PR数だけで生産性や品質を決めないことが重要です。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>arXiv論文：AIの道具操作はJSONよりコードが有利な場合——14モデル中11、長い連鎖で差18.8ポイント</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-programmatic-tool-calling/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-programmatic-tool-calling/</guid>
    <description>BFCL v4でコード型とJSON型のツール呼び出しを比較。新しいモデルほどコード型が強い一方、旧モデルには構文エラーもあった。 複数ツールの連鎖や大量並列ではコード型が有力ですが、安全な隔離実行とモデル別の構文テストが欠かせません。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>The Bitter Lesson of Tool Calling（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：正しいAIも誤ったヒントで答えを変更——1000問×4条件、全モデルで弱点を確認</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-mist-selective-trust/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-mist-selective-trust/</guid>
    <description>同じ問題に正しい・誤り・無関係の文脈を加えるMISTを公開。何でも無視せず必要な情報だけ信じる学習法SCOPEも提案した。 検索結果や資料に誤りが混ざると強いAIも答えを変えます。重要判断では引用元を開いて根拠を確認してください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>MIST・SCOPE（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AI向け端末課題5431件を難易度調整——強いAIだけ解ける条件で学習効果を向上</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-calibforge-terminal-tasks/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-calibforge-terminal-tasks/</guid>
    <description>作れるかだけでなく、複数AIの合否を見て課題を書き直すCalibForge。端末操作と別のコード評価で改善を確認した。 AI教材は量だけでなく、解ける強さの境界と自動採点を確認することが、実務課題への転用で重要だと分かります。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>CalibForge（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AI対戦の強さを中央値で74分の1の試合数で判定——途中終了しても信頼度を保証</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-av-aivat/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-av-aivat/</guid>
    <description>運の影響が大きいゲームでも、証拠が十分になった時点で安全に評価を止めるAV-AIVATを提案。 AI同士の比較を早く安く終える設計の手掛かりですが、74倍という数字を別の課題へそのまま当てはめないでください。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>研究論文</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AIは自分の作業環境をどこまで改善できるか——5モデル・4課題・111回で評価</title>
    <link>https://ai.manapick.app/news/n-2026-08-08-arxiv-harnessopt-bench/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-08-arxiv-harnessopt-bench/</guid>
    <description>プロンプトや道具、記憶、制御手順をAI自身に直させる能力を測るHarnessOpt-Benchを提案。 AIの性能はモデル名だけでなく周囲の設定でも変わります。比較時は同じ予算と見えないテストで確かめる視点が重要です。</description>
    <pubDate>Fri, 07 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>研究論文</category>
    <category>search</category>
  </item>
  <item>
    <title>Google、台風AI「WeatherNext 2」をオープン化——3日先予報が従来の2日先と同等、最大1000通りを計算</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-weathernext-cyclone-open/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-weathernext-cyclone-open/</guid>
    <description>台風の進路・強さ・風の広がりを1つのAIで予測。Nature掲載研究に合わせ、コードと重み、無料Colab版を公開した。 台風予測を地域向けに検証できる公開AIが増えますが、避難判断は必ず気象庁などの公式情報を優先してください。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>WeatherNext 2（Google DeepMind）</category>
    <category>search</category>
  </item>
  <item>
    <title>Mistral、画像も文章も安全判定する「Shieldstral」公開——30億規模、16GB GPUで動作</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-mistral-shieldstral/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-mistral-shieldstral/</guid>
    <description>自然文で安全基準を渡せる30億パラメータの判定AI。Apache 2.0の重みを公開し、再学習なしで用途別の基準に対応する。 手元のGPUで用途別の安全判定を試せますが、日本語や誤判定を検証し、人の確認を組み合わせる必要があります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Mistral Shieldstral</category>
    <category>local</category>
  </item>
  <item>
    <title>arXiv論文：AIの推論教材を50種類の生成器で作成——正解可能でも学習に役立つとは限らない</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-reasoning-core/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-reasoning-core/</guid>
    <description>数学・論理・計画・コードなど50種類の手続き型教材を公開。30億規模の比較で3評価の平均が既存教材を上回った。 推論AIの学習データは件数だけで選ばず、答えの形式、難しさ、採点の正しさを監査する重要性が分かります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Reasoning Core（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：複数リポジトリをまたぐコード教材「OctoLong」——長文教材の12%置換で理解力が向上</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-octolong-code-context/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-octolong-code-context/</guid>
    <description>AST解析・言語サーバー・パッケージ管理で依存関係を追い、数百万トークン級のコード教材を構築。18モデルと比較した。 コードAIへ長文を与えるだけでなく、実際の依存関係をたどる教材設計が重要だと分かります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>OctoLong（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：仕事AIの記憶を1005課題で検証——短い要約より経験豊富な記録が有効、誤記憶には弱い</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-contextweave-memory/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-contextweave-memory/</guid>
    <description>14人の数カ月分の仕事を匿名化して再構成。強い記憶方式で作業点は68.08から78.20、好み一致は41.50から70.60へ上昇した。 仕事AIの記憶は詳しいほど役立つ可能性がありますが、出典・更新日・訂正手段を一緒に設計する必要があります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>ContextWeave（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>Google、AIエージェント部品を1つに束ねる「Agent Plugins 1.0.0」を支持——SkillsとMCPを共通形式で持ち運び</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-agent-plugins-standard/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-agent-plugins-standard/</guid>
    <description>GoogleがAgent Pluginsの中核運営に参加。plugin.jsonと固定フォルダで、SkillsやMCPサーバーを複数のAI開発環境へ配れる。 AI向けの手順書とMCPを複数環境で再利用しやすくなりますが、導入時の権限・安全確認は利用者側の責任です。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Agent Plugins 1.0.0</category>
    <category>code</category>
  </item>
  <item>
    <title>AWS、AIの「行動履歴」で許可を決めるTemporal Policiesを公開——高額操作は毎回人の承認</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-aws-agentcore-temporal-policies/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-aws-agentcore-temporal-policies/</guid>
    <description>AgentCore Gatewayを通るAIの操作履歴を見て、順序、データ鮮度、累計額、人の承認を外側から強制する仕組み。 AIに重要操作を任せる際、過去の行動まで含む規則で止められますが、Gateway外の操作や設計漏れは別に対策が必要です。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Amazon Bedrock AgentCore</category>
    <category>code</category>
  </item>
  <item>
    <title>GitHub CopilotにKimi K3追加——ただしActions障害対応で展開を一時停止、入力100万トークン3ドル</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-github-copilot-kimi-k3-paused/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-github-copilot-kimi-k3-paused/</guid>
    <description>Kimi K3をCopilotの主要画面へ段階展開する計画。GitHub Actionsの問題を受け一時停止中で、組織版は初期状態で無効。 CopilotでKimi K3を選べる予定ですが、現在は展開停止中です。再開確認と組織管理者の許可を待つ必要があります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot / Kimi K3</category>
    <category>code</category>
  </item>
  <item>
    <title>arXiv論文：画像を切り抜くAIは「見たふり」も——6モデル・5評価で、画像が回答に効かない失敗を確認</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-visual-tool-use-audit/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-visual-tool-use-audit/</guid>
    <description>画像AIの切り抜き・拡大操作を因果的に検査。返された画像が回答に影響しない型と、見る順序が不適切な型を分けた。 画像AIが道具を何回使ったかではなく、得た画像が回答の根拠になったかを確認する重要性が示されています。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>arXiv研究（画像ツール利用の因果監査）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：0.48MBの手話認識AIをスマホで実行——1万874枚を専門家確認、1画像3.98ミリ秒</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-bangla-sign-mobile/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-bangla-sign-mobile/</guid>
    <description>ベンガル手話38種類の実写データと約30万パラメータの小型モデルを公開。一般的なスマホで高速動作した。 手話認識を端末内で動かす現実的な小型化例ですが、日本手話や連続会話への適用には別データと検証が必要です。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>arXiv研究（ベンガル手話認識）</category>
    <category>search</category>
  </item>
  <item>
    <title>ACM Multimedia 2026論文：混雑動画の数え間違いを62.01%削減——奥行きと重複除外を組み合わせ</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-acmmm-depth-video-counting/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-acmmm-depth-video-counting/</guid>
    <description>文字や画像で指定した物体を混雑動画から数えるDG-Det。RGBと奥行き、遮蔽予測、フレーム間の重複除外を統合した。 混雑映像の数え上げが改善する可能性がありますが、62.01%は特定評価での誤差減少で、現場ごとの再検証が必要です。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>ACM Multimedia 2026研究（動画物体数え上げ）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：動画AIは速い出来事の数え上げに弱い——2190本で検証、高頻度では正答0.2%</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-video-event-counting/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-video-event-counting/</guid>
    <description>出来事の回数と頻度を分けて検査。高回数・高頻度では正答が0.2%、実際の出来事を拾えた割合は18.1%だった。 動画AIの自然な説明と、出来事の回数・時刻の正確さは別です。重要な集計は元動画や時刻記録で確認する必要があります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>動画言語モデルの時間理解（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：動画生成AIは見た目が自然でも物理量を誤る——22課題でシミュレーターと比較</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-gauge-physics-video/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-gauge-physics-video/</guid>
    <description>現実の測定値に基づくGAUGEを提案。3つの物理エンジンと6つの画像動画AIを調べ、万能な方式はないと報告した。 ロボットや設計に生成動画を使う場合、自然な見た目だけでは不十分です。加速度や衝突など測れる物理量で検証する必要があります。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GAUGE（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：調査AIの確認質問を約3分の1に——目的と根拠を整理して依頼文を改善</title>
    <link>https://ai.manapick.app/news/n-2026-08-07-arxiv-gsteer-deep-research/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-07-arxiv-gsteer-deep-research/</guid>
    <description>G-STEERは利用者の目的・制約・好みを依頼文へ反映。2種類の調査AIで、比較手法より少ない質問で個人化を高めた。 調査AIへ予算・期限・重視点を先に整理して渡すと、聞き返しを減らしながら自分向けの報告に近づけられます。</description>
    <pubDate>Thu, 06 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>G-STEER（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>Google Classroom、Geminiで課題ごとの評価表を直接作成——先生が確認・編集して追加</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-gemini-classroom-rubric/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-gemini-classroom-rubric/</guid>
    <description>課題作成画面の内容を基に、採点基準となるルーブリックをGeminiが下書き。Education各エディションへ8月5日から展開する。 先生は評価表のたたき台を短時間で作れますが、採点の公平性を守るため、基準と配点の人による確認が欠かせません。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Gemini</category>
    <category>work</category>
  </item>
  <item>
    <title>arXiv論文：AIがコンパイラの見逃した高速化を発見——最良モデルは93.3%の課題で性能向上に成功</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-segabench-compiler/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-segabench-compiler/</guid>
    <description>C/C++の最適化を120課題で検証。最良モデルは正しい成果物を94.8%の回答で作り、実行速度の改善も多くの課題で確認された。 AIによるコード高速化は有望ですが、提案をそのまま採用せず、テストと実測を通すことで安全な補助役として使えます。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>SeGaBench（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AIエージェントの「スキル蓄積」は常に効くとは限らない——5分野500課題で検証</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-continual-skill-bench/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-continual-skill-bench/</guid>
    <description>順番に課題を解くと成績は上がる一方、明示的なスキル保存は平均で会話内学習と同程度。再利用できる手順の整理に課題が残った。 エージェントの記憶やスキル庫は、保存件数ではなく、別の仕事へ再利用できたかを試す必要があると分かります。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>ContinualSkillBench（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AIの「考える量」を同じ予算だけで比べるのは危険——推論法を3種類に整理</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-test-time-scaling/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-test-time-scaling/</guid>
    <description>1回答を長く考える、複数回答を選ぶ、途中状態を探索する方式を区別。20億件超の推論記録も公開用に整備した。 AIの比較表では、点数だけでなく、何回回答させ、どう選び、どれだけ計算したかを見ないと、費用対効果を誤解します。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>推論AIの評価方法（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>Google、3種類の生成AIを1つの入口で切り替え——モデル振り分けを公開プレビュー</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-google-api-gateway-model-routing/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-google-api-gateway-model-routing/</guid>
    <description>OpenAI互換の入口から、同じVertex AIホスト上の複数モデルへ動的に振り分ける機能を公開プレビュー。独自プロキシの管理を減らせる。 複数モデルを試すための接続コードをまとめやすくなりますが、同一Vertex AIホスト内という制約とプレビュー段階である点に注意が必要です。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Google Cloud API Gateway</category>
    <category>code</category>
  </item>
  <item>
    <title>arXiv論文：40億規模の検索AIが約300億規模に匹敵——8,500例と「答えから逆算する採点」で学習</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-abseeker-search/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-abseeker-search/</guid>
    <description>検索の途中手順を答えから逆算して細かく採点するABCを提案。ABSeekerはBrowseCompで37.3%、文脈管理ありで55.3%を報告した。 検索AIを大きくするだけでなく、途中の良い行動を細かく教える方法が有効なら、費用を抑えた深掘り調査につながります。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>ABSeeker（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AI安全性テストを約10問まで圧縮できる場合——192モデル・8評価を項目反応理論で分析</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-safety-irt/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-safety-irt/</guid>
    <description>安全性評価を項目反応理論で整理。選んだ約10問で一部ベンチマークを近似し、評価費用を97〜99%削減できると報告した。 安全性評価の費用を下げつつ、平均点の裏にある弱点を見つけやすくなる可能性がありますが、少数問題だけで安全を断定してはいけません。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>IRT for AI Safety（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：科学コードAIの低得点はテスト側の欠陥が原因——65課題を修正すると正答率84〜98%</title>
    <link>https://ai.manapick.app/news/n-2026-08-06-arxiv-scicode-verified/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-06-arxiv-scicode-verified/</guid>
    <description>SciCode全65課題を専門家が監査し263件の欠陥を報告。修正版では12モデルの小課題正答率が45〜60%から84〜98%へ上昇した。 AIの順位や進歩を判断するとき、モデルの点数だけでなく、評価問題そのものが正しく作られているか確認する大切さが分かります。</description>
    <pubDate>Wed, 05 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>SciCode-Verified（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AIエージェントの失敗を約200マイクロ秒で検知・再実行——成功率52%から73%へ</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-agent-failure-repair/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-agent-failure-repair/</guid>
    <description>2,823回の実行記録から軽量監視と決定的検証を比較。失敗を検知して巻き戻す仕組みが、単純な再試行より高い回復率を示した。 AIの「完了しました」を信じるだけでなく、道具の結果と必要操作を自動照合すると、低コストで失敗を発見・回復できる可能性があります。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>AIエージェント障害検知（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>AIMLSystems 2026論文：AI評価は課題の15〜25%で結論できる場合も——途中点だけの公表に警鐘</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-aimlsystems-parevallayer/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-aimlsystems-parevallayer/</guid>
    <description>2つのAIを比較する途中結果に、事前の判定規則と保留を追加。3つの公開評価では全課題の15〜25%で最終判断と一致した。 AIモデルの比較費用を減らせる可能性がある一方、途中点だけでは誤解を招くため、停止規則と未判定数の公開が重要です。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>ParEvalLayer（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：深掘り調査AIを500課題・31分野で検証——根拠確認まで追える評価データを公開</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-deep-research-benchmark/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-deep-research-benchmark/</guid>
    <description>簡単な質問を多段階の調査課題へ自動発展させ、各手順と確認点をDAGで記録。10大カテゴリ・3形式を収録した。 深掘り調査AIを、文章の見栄えではなく、必要な手順と根拠を満たしたかで比較できる公開基盤です。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>深掘り調査ベンチマーク（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：学校向けAIの安全性を10モデルで検証——教育特有の危険と長い会話に弱さ</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-eduzone-k12-safety/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-eduzone-k12-safety/</guid>
    <description>児童生徒・教員の利用場面に6種類・28小分類の危険を組み合わせ、1回の質問から動的な長い会話まで評価した。 教育AIの安全確認では、一般的な禁止事項だけでなく、授業内容・利用者の立場・長い会話を組み合わせた試験が必要だと示します。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>EduZone（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AI本体を変えず「実行の仕組み」を失敗から修正——成功率44.3%から53.6%へ</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-harness-r1/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-harness-r1/</guid>
    <description>別の9Bモデルが失敗記録から検証済みパッチを作成。WebShopなど3評価で、Qwen 3.5 9Bの課題成功率を9.3ポイント改善した。 AI本体の再学習が難しくても、失敗ログから文脈・道具・検証・回復の仕組みを直す改善方法が考えられます。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Harness-R1（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：データセンターの運用ルールをAIが設計——配置・増減・電力の3課題で専門家案を上回る</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-atumai-datacenter/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-atumai-datacenter/</guid>
    <description>自然文の目標を機械検査できる仕様へ変換し、複数の探索法で候補を提案・試験・改善するAtumAIを発表した。 AIに業務ルールを作らせるとき、自然文だけでなく、守る条件と評価方法を機械検査できる形にする重要性が分かります。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>AtumAI（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>GitHub Copilotのクラウドエージェント、課題ごとに「考える量」を選択可能——高設定はAIクレジット増に注意</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-github-copilot-reasoning-level/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-github-copilot-reasoning-level/</guid>
    <description>対応モデルの推論レベルを、クラウドエージェントへ仕事を渡す時に指定できる。複雑さと消費クレジットのバランスを調整しやすくなった。 難しい課題にだけ多く考えさせ、簡単な作業ではクレジット消費を抑える、といった使い分けがしやすくなります。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>GitHub Copilot、自動化をIssue・PRコメントから起動——文書更新、エラー調査、後続課題の作成に対応</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-github-copilot-comment-automations/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-github-copilot-comment-automations/</guid>
    <description>あらかじめ決めたコメント文を合図にクラウドエージェントを実行。レビュー中の依頼を、その場から繰り返し自動化できる。 PRやIssueで会話している場所から定型のAI作業を始められ、文書更新やエラー調査の依頼手順をチームで統一できます。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>GitHub Copilot</category>
    <category>code</category>
  </item>
  <item>
    <title>Google ClassroomのGemini、年齢制限を拡大——8月10日から小中高生も授業資料でクイズ・暗記カードを作成</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-gemini-classroom-all-ages/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-gemini-classroom-all-ages/</guid>
    <description>管理者が利用を許可した全学齢の生徒へGeminiタブを拡大。クラスや課題を選ぶと、授業内容に沿った学習支援を受けられる。 学校が許可すれば、小中高生も授業資料に基づくクイズや学習ガイドをClassroom内で作れます。AIの答えを教材と照合する習慣も重要です。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Gemini</category>
    <category>work</category>
  </item>
  <item>
    <title>Google MeetのAI議事録、共有画面のスクリーンショットを自動追加——グラフや図の説明不足を補完</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-google-meet-notes-screenshots/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-google-meet-notes-screenshots/</guid>
    <description>「Take notes for me」が発表中の画面を議事録へ保存。文字起こしだけでは伝わりにくい図表の文脈を残せる。 議事録だけ読んでも図表の意味を追いやすくなります。一方、共有画面が文書に残るため、機密情報を映す会議では設定確認が欠かせません。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Gemini in Google Meet</category>
    <category>work</category>
  </item>
  <item>
    <title>arXiv論文：検索できる最先端AIのW杯勝敗予測は平均63.9%——ブックメーカー本命と同水準、AI多数決も効果なし</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-worldcup-arena/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-worldcup-arena/</guid>
    <description>2026年W杯の全104試合を試合前に予測し、答え漏れのない実時間評価を実施。6モデルは互いに似た予想へ偏った。 AIを複数集めても、同じ本命へ偏れば予測は強くなりません。未来の判断では説明の詳しさより、基準となる単純予測との比較が重要です。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>AI予測評価（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>arXiv論文：AIへの音声入力は文字入力より回答精度を下げやすい——考える量を増やしても音声の弱さは残る</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-arxiv-voice-keyboard-robustness/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-arxiv-voice-keyboard-robustness/</guid>
    <description>音声認識の言い直しやキーボードの打ち間違いを再現して比較。音声では余計な語より、文の組み替えによる情報欠落が影響した。 音声入力では、AIの推論力より前に文字起こしで条件が消えることがあります。送信前に文字を確認するだけでも重要な誤解を減らせます。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>音声・文字入力のAI耐性（AI研究）</category>
    <category>search</category>
  </item>
  <item>
    <title>AWSだけで“本家Claude”を契約——Claude Platform on AWS一般提供、ただしデータはAWS境界外</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-claude-platform-on-aws/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-claude-platform-on-aws/</guid>
    <description>AnthropicのネイティブAPIとClaude Code・CoworkをAWSの認証、請求、監査で利用可能に。Amazon Bedrockとはデータ処理場所が異なる。 AWSの認証・請求を保ったままClaude本体の新機能を使えます。一方、Bedrockと同じデータ境界ではないため、企業導入では違いの確認が欠かせません。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Claude Platform on AWS</category>
    <category>work</category>
  </item>
  <item>
    <title>Genspark、AI版Officeを丸ごと公開——Word・Excel・PowerPoint・PDFを1つに、ソースコードもApache 2.0</title>
    <link>https://ai.manapick.app/news/n-2026-08-05-genspark-genoffice-open-source/</link>
    <guid isPermaLink="true">https://ai.manapick.app/news/n-2026-08-05-genspark-genoffice-open-source/</guid>
    <description>macOS・Windows向けの文書、表計算、スライド、PDF編集アプリを公開。AIが会話だけでなくファイルを直接編集し、差分も確認できる。 Office形式を編集できるAIアプリの中身まで確認・改良できます。ただしAI処理は完全ローカルではなく、Gensparkのアカウントとクレジットが必要です。</description>
    <pubDate>Tue, 04 Aug 2026 15:00:00 GMT</pubDate>
    <category>AIニュース</category>
    <category>Genspark GenOffice</category>
    <category>work</category>
  </item>
</channel>
</rss>
