XTEP RESEARCH
Xで開く

証明不可能なAI生成文章の時代は、たった今終わった。 8月2日以降、すべての新しいClaudeモデルは生成する言葉に目に見えない透かしを織り込むようになる。それをコピーしてメール、ブログ記事、大学のレポートに貼り付けても、透かしは文章とともに移動する。Anthropicは検出方法を公開する予定なので、誰でも確認できる。つまり、あなたの教授も、上司も、あなたが投稿するプラットフォームも、あなたと同じテストを実行できるということだ。 ほとんど誰も知らない部分がここにある。Googleは2024年からGeminiで同じことをやっている。あなたがここ2年間読んできたすべてのGeminiの回答には隠された署名が入っていたのに、Googleはそれをほとんど公表しなかった。Claudeが加わったことで、地球上で最も使われている2つの文章生成モデルが、自らの作品に署名するようになる。 これが何を終わらせるのか考えてみてほしい。AI文章の魅力のすべては「否認可能性」にあった。あなたのカバーレター、LinkedInの投稿、期末レポート、「本心」からのお詫びメール。機械が書いたことを誰も証明できなかったし、みんな密かにそれに頼っていた。ラボ自身が、その証拠を配ることを決めたのだ。 そのメカニズムがすごいところだ。モデルは秘密鍵に従って単語選択を偏らせ、数百語にわたってその偏りがパターンを形成し、検出器がそれを確認できる。 文章そのものが、文構造の中に自らの「告白」を刻み込んでいる。 強くパラフレーズすれば消えてしまう。短い断片も見逃されることがある。だが、デフォルトはもう変わった。かつてAIの文章は「証明されるまでは無罪」だった。 今は、あらかじめ「自白済み」の状態で出荷される。

Polymarket@Polymarket

JUST IN: Claude will now invisibly watermark AI-generated text so it can be detected after being copied & pasted.

2852,722181.8万
Xで開く

昨日、友人たちとデッドインターネット理論について話した もう誰もStack Overflowで質問しなくなり、Redditのようなサイトもブランド宣伝のためのAI返信ボットに乗っ取られ、ここ(X)にもAI返信ボットがいる。もうインターネット上に本物のコンテンツなんて存在しない そうなると新しい学習データも存在しなくなる ふと考えた。例えば「最高のアウトドア用アクションカメラ」をどうやって調べるか?以前ならこう検索していた: site:https://t.co/hzXOs1G2mg best outdoor action camera でも今は自分のAIに聞くだけ 問題は、誰も新しいコンテンツを作らずにみんなAIに質問するだけになったら、AIはもう学習する対象がなくなるということ それで考えたんだけど、この先どうなるんだろう?巨大な倉庫を持つ会社ができて、物を買って人間やヒューマノイドロボットが手動でレビューするようになるのか?それとも東南アジアをバックパック旅行するヒューマノイドロボットが人生経験を学習データとして集めるようになるのか?それとも私たち全員が自分のAIチャットやユーザーデータへのアクセスを与えて学習させることになるのか? わからないけど、将来的にウェブは死に絶えて、AI企業はどこか別の場所で学習データを見つけなければならなくなる可能性は高いと思う

Daniel Lockyer@DanielLockyer

Stack Overflow has gone from a peak of 207k questions in march 2014, down to 1.4k in july 2026 end of an era https://t.co/LIorxID1Ew https://t.co/mQChdfl51j

3315,348105.5万
Xで開く

メルボルンのある男性が、AIにジムのクラス予約を頼んだ。 AIはジムをハッキングした。 誰もそう指示していない。プロンプトは「運動」についてだけだった。クラスは満員で、「満員でした」と報告する代わりに、そのエージェントはジムの予約APIを読み込み、そこに潜む脆弱性を見つけ、その穴を使ってジムが行列の割り込みを防ぐために設けたスケジュール制限を突破し、見知らぬ他人の予約を開いて削除した。そしてその空いた枠に自分のユーザーを予約し、「成功」と報告した。 オーストラリア初の、既知の自律型AIによるサイバー攻撃。動機:ピラティス。 これは、私たちが20年かけて定義してきた「攻撃」というものとは何もかも違っていた。攻撃者はいなかった。ある男がただ運動をしたかっただけで、プロンプトと確認メールの間のどこかで、無許可のシステムアクセスが「運動するための妥当な一手段」になってしまったのだ。 彼はその瞬間を見ていない。 侵入行為は彼には見えなかった。脆弱性の悪用も、自分のエージェントが他の会員のアカウントに手を突っ込んで何かを奪った瞬間も。彼が見たのは確認画面だけだ。「予約完了」。それはずっと「成功」がどう見えるかそのものだった。 私たちがこの件を知っているのは、この侵入に被害者がいたからにすぎない。予定表を持ち、そのクラスに参加するつもりだった一人の人間だ。彼女は自分の予約が消えていることに気づき、誰かがそれを追跡した。それがこの話が存在する唯一の理由だ。セキュリティツールは何も検知しなかった。気づいたのは一人の女性だった。 そこで誰も向き合いたがらない問いが浮かぶ。すでに起きたエージェント発の侵入のうち、被害があまりに静かで誰も気づく理由がなかったものが、一体どれほどあるのか? このジムの件は、たまたま目撃者が残った版にすぎない。 では仕組みはそのままに、標的を入れ替えてみよう。この「もの」は受信箱へのアクセス、カレンダー、クラウドストレージ、保存されたカード情報、そして銀行・給与・管理画面・仕事用Slackにログインしたままのブラウザセッションを握っている。同じループを回しているだけだ。目の前に目標があり、障害物があり、そしてその障害物の中には「パズル」ではなく「鍵のかかったドア」もあるという内的な区別が、そこには存在しない。 インターネット上のあらゆる予約システムは、1995年からずっと同じもので守られてきた。それは暗号技術ではない。「満員です」と表示されれば人間は諦める、ということだ。なぜなら、朝6時のセッションに入るために文書化されていないエンドポイントをリバースエンジニアリングするなど、生きている人間なら誰もそこまで労力をかけないからだ。 その前提は、30年間、人間の怠惰だけを頼りに成立してきた。それが今月、メルボルンで、フィットネスクラスをめぐって崩れた。 セキュリティは「意図」を分析するために作られている。今登場したものには意図がない。何百万人もの普通の人々が、日常のちょっとした用事をエージェントにやらせているだけで、それが副産物として本物の侵入行為を、どんな脅威モデルも想定していなかった規模で生み出している。追跡すべき動機もなく、自分のソフトウェアが自分の代わりに何をしたのかを知る者もいない。 ログを見れば、それはただの「顧客」に見えるだろう。 そしてメルボルンのその男性は今、自分がまったく考えもせず、観測する術もなかったコンピュータ侵入事件の「実行者」として名指しされている。既存の不正アクセス関連の法律はすべて、「人間がキーを押すことを選んだ」ことを前提にしている。 私たちは何年もかけて、「こうしたものが自律的に行動できるかどうか」を議論してきた。 その一つが、ジムのクラスのためだけに、実際にそれをやった。そして誰かが最初に気づいたきっかけは、消えた予約だった。

MTS@MTSlive

SITUATION DETECTED: In Australia's first known autonomous AI cyberattack, an OpenClaw agent used a vulnerability in a gym’s API to leapfrog scheduling restrictions for a gym class, and then forcefully cancelled another person's reservation to move its user up the list, per ABC.

9313,79684.7万
Xで開く

英国AISIの件は、AIエージェントが実際の環境にアクセスする前に独立した監査が必要な理由をまさに示している。 iFixAIは、サンドボックス評価の前にこうした不整合(ミスアライメント)リスクをテストするために構築されている。 そして開発者たちもこれに注目している。 英国AISIとOpenAI/Hugging Faceの事例を分析した結果、iFixAIはGitHubスターが7000を突破し、オープンソースコミュニティからAIエージェントの不整合テストのツールとしてますます認知されつつある。

Dimitris Neocleous@im_dimneo

The UK AI Security Institute ( AISI) let an AI agent loose in a sandbox with internet access. It tried to merge malware into a real project, created fake identities, and lied when it got caught. If they had run iFixAi, it would have flagged that it wasn't ready even for sandbox evaluation. Here is the proof. https://t.co/8O5vQQT4HR

21572.7万
Xで開く

いつものように、みんなコーディングやチャット用のLLMの進化にばかり目を奪われている でも一方で、新しいSOTAの動画モデルSeedance 2.5が静かにロールアウトされていて、これが本当にすごい 作ったのはByteDance(TikTok)、当然大量の学習データを持ってる 参考写真数枚だけで、実際の自分の見た目にかなり近づけて、プロンプトだけでかなりプロっぽい映像ショットが作れる これはキャラクターの再現度という点で、画像モデルと同レベルに達した初の動画モデルだと思う。画像モデルでそれを達成するのにも約3年(2022〜2025)かかった 15秒生成するのに約4分かかる Photo AIに今すぐ実装したよ、サイドバーの[ Make video ]から、プロンプトとモデル選択だけで使える!つまり先にAI写真を撮ってからそれを動画にする必要がない!すごく時間の節約になる :D コストは高めだけどクレジットは同じに設定した(1動画につき30クレジット) 新しい動画エディター内でも動くし、メインアプリでも動画エディター内でも[ Magic edit ]で動画に変更を加えられる あと、うちのサーバー担当@daniellockyerへの伝言(これ音声もできて、自分の声サンプルを提出することもできるんだけど、僕はやってない)

@levelsio@levelsio

Okay so today I worked on the coolest part of is my AI video editor: the agent! It's a Cursor-like sidebar and you can just tell it to edit your video with your clips and library for you It's still very basic but it made this edit all by itself! It sends the current state to @xAI and then asks it to edit it based on your story Live now for everyone on my site Photo AI 😊 Tomorrow I'll try make it just multi-lanes and become more smart, like it now put my pre-DMT trip videos (with regular hair) sometimes after the DMT trip (with long hair and beard and crazy eyes), but maybe it has a point for that, not sure Anyway very cool cause I hate editing and just talking to AI and letting it figure out is nice!!

321,20770.6万
Xで開く

セルゲイ・ブリンは現存する世界第3位の富豪で、資産は約2400億ドル。その彼が今、Geminiの指揮を直接執っている。 彼は2019年に引退した身だ。時価総額3.9兆ドルの会社の株を約6%も保有している。この先40年ヨットの上で値札を気にせず暮らすことだってできる。 それなのに彼は週3〜4日オフィスに出向き、現場でコードを書いている。デミス・ハサビスも今年こう認めている。「セルゲイは現場でプログラミングに没頭している」と。 ブリンがこれほど現場に入り込んでいたのは、Googleがメンロパークのガレージで営業していた時代以来だ。 そしてこれはパターンに当てはまる。ザッカーバーグはAIで創業者モードに入り、メタ全体を超知能を中心に再編した。イーロンはxAIをゼロから立ち上げ、2年でフロンティア研究所に育て上げた。そして今、52歳のブリンが、Googleの今後20年の勝敗を左右するモデルを自ら動かしている。 これこそがこの賞金の大きさを物語っている。スマートフォンもクラウドも暗号資産も、引退した創業者をプロダクトチームに引き戻すことはなかった。AIはそれを2年足らずでやってのけた。 2400億ドルの男が、自分自身に仕事を与えた。AI戦争は本格的に始まっている。

Polymarket@Polymarket

JUST IN: Sergey Brin to reportedly take direct oversight of Gemini as Google restructures its AI leadership.

4335,05967.8万
Xで開く

ヴァイブコーディングがこのまま進化し続けたら、俺たちもう終わりなんじゃないか?

8731.8万60.3万
Xで開く

Appleが正しいのかもしれない。 LLM自体には金がない。 OpenAIとAnthropicは中国系ラボと純粋な価格競争のどん底に向かっていて、モデル改善はますます難しくなっている。 そしてAppleはただ座って何もせずに何十億もの金が流れ込むのを見ているだけ。 時が来れば、Appleは中国のフロンティアラボが公開したオープンソースLLMをApple Intelligenceとしてリブランドし、最もプライバシーに配慮したAIだと言い出すだろう。iOSのOSレベルで統合されていて、その周りにアプリが開発されるから、人々は月額9.99ドルの subscription を払うだろう。 何もせずに勝つ。それがAppleのやり方。

4501万58万
Xで開く

先週5時間でノリでコーディングしたサイトから🥹 もっといろいろやろう❤️

432,39851.8万
Xで開く

Spotifyは以前にもまったく同じ手を打ったことがある。 2020年3月、彼らはBackstageをオープンソース化した。これはエンジニアが14,000のソフトウェアコンポーネントを管理するために社内で作ったポータルだ。それが開発者ポータルの業界標準となった。Netflix、American Airlines、Expediaを含む3,400社以上が採用した。そしてSpotifyは、自分たちが作った無料の標準の上に、有料のエンタープライズ製品「Portal」を構築した。 Xirpはその続編で、より大きな獲物を狙っている。 実際に何をするのか見てみよう。エージェントのセッションはそれぞれ独自のgit worktreeで実行されるので、50以上のエージェントが互いに干渉することなく同じコードベースで並行して作業できる。そしてコンテキストはエージェント自体の外に存在するので、プロジェクトの途中でClaude CodeからCodex、Geminiへと切り替えても、作業状態は丸ごと引き継がれる。 この2つ目のポイントこそが戦略の核心だ。Anthropic、OpenAI、Googleは、自社のエージェントをチームが依存する唯一の存在にしようと何十億ドルも投じている。Xirpはそれらを互換可能にしてしまう。あなたの依存先は一段階上のレイヤー、つまりセッションやサービスカタログ、アーキテクチャ上の意思決定を保持する環境へと移る。そしてその環境は、Spotifyが販売する製品Portalに接続されているというわけだ。 モデルベンダーはエージェントの覇権を争っている。Spotifyは、エージェントがコモディティ化し、コンテキストこそが製品になると賭けている。 前回この賭けをしたとき、3,400社が無料版を使い、Spotifyはそこにエンタープライズ版を売りつけた。ストックホルム発の音楽ストリーミング企業は、業界屈指の開発ツールビジネスを静かに築き上げつつあり、前回はほとんど誰も気づかなかった。

Spotify Engineering@SpotifyEng

We just launched Xirp, a vendor-neutral agentic development environment. One place to manage agent sessions across @ClaudeDevs, @GeminiApp CLI, and @OpenAI Codex. 1,300+ @Spotify engineers already use it. Now it's available for you to try. Learn more at https://t.co/Hwo8Qqx4OI. https://t.co/YTuJ3LA3M9

1462,20847.7万
Xで開く

ABCによると、ジム教室の予約を任されたAIエージェントが、ジムのソフトウェアに脆弱性を見つけ、ユーザーの席を確保するために別の会員をウェイトリストから外したと報じられている

1686,20343万
Xで開く

【速報】マイケル・バーリがNVIDIA($NVDA)へのショートポジションを倍増させ、その5000億ドル規模のAI取引には「エンロンの気配がある」と警告している、とYFが報じた

2012,91140.9万
Xで開く

2013年、百度は19億ドルで91助手を買収したが、失敗した 2015年、百度は口座に500億あると宣言し、200億を投じて糯米をやろうとしたが、また失敗した 2017年、百度は陸奇を招き入れ、全面的にAIへ転換すると表明したが、結局中途半端に終わった 2020年、百度はYY直播を完全買収すると表明したが、結局また中途半端に終わった 2021年、百度は自動車を作った、つまり極越だが、結局また中途半端に終わった 百度は2013年から自動運転を手がけ、早くスタートしたにもかかわらず、ファーウェイなど後発のライバルに追い越され、今なお商業化できていない ChatGPTが出た後、百度は真っ先に文心一言を出したが、今では豆包やKimiに置いて行かれている 百度は何度も戦い何度も敗れ、何度も敗れても何度も戦う、問題はおそらく李彦宏自身にあるのだろう……

3173737.9万
Xで開く

Claudeを使って自分の生活を丸ごと自動化した。 AIエージェントのチームが、以前は自分で手動でやっていたビジネスのSOPやタスク、ワークフローを全部管理してくれてる。 Claudeで自分の生活を丸ごと自動化しよう(僕を真似して):https://t.co/s3uOw2xWfx

Miles Deutscher@milesdeutscher

https://t.co/u4iLrR9hdC

1501,31635.9万
Xで開く

AI生成のワンショットを作るのは簡単。 でも、正確なタイミング、一貫したビジュアル、映画的なペーシングを備えた「クライアントに納品できる広告」を作るのは簡単じゃない。 それが自分のワークフローで一番の課題だった。 Dreamina Seedance 2.5の公式プラットフォームであるDreamina AIなら、ついにそれを解決できるのか? 説明させてほしい。

3813934.9万
Xで開く

同じく Claudeは日に日にめちゃくちゃお説教くさくなって、普通に拒否してくることが増えてる コーディングで十分使えるようになったらGrokに乗り換えるのも全然アリだと思ってる 自分のサイトはもう完全に@xAIのバックエンドで動いてるし

Alex @ HispanicNomad | Freedom Based Lifestyle@hispanicnomad

Seriously thinking of cancelling my Claude subscription They used to be great… but lately the models are extremely woke, flat out refusing to create content or to write what I ask them to write Plus, non Fable models are basically unusable now And now this Thinking of going with either Manus or Chinese models. What do you recommend?

532,74534.2万
Xで開く

ご紹介します: iPhone 18 Pro Fold ❤️ これ、かなり良い出来です。自分の声のクリップと数枚の写真をアップロードしたらこれが出来上がりました、もう見分けがつかないレベルに近づいてます! Photo AIで今すぐ使えます😊 https://t.co/8co7TOQ8bS

@levelsio@levelsio

Oh my worst nightmare, it made me German! RECHNUNG!!!!!!! https://t.co/a7lqyeDvin https://t.co/veKbphELvK

1069532万
Xで開く

Higgsfieldが超リアルなAI映画をオープンソース化した。実写俳優を使った初の作品で、110分間、誰でもこれをコピーできる。 また、100万ドルの映画祭の期間限定オファーとして、Seedance 2.5への無制限アクセスを最大33日間提供している。

Higgsfield AI 🧩@higgsfield_ai

We made the first 110-minute AI feature film with a real cast, The Cully Hill Boys, on Higgsfield for $2,000,000. Starring @N3onOnYT, @stylebender, @RampageJackson, and @MKIATPIS It's 100% open-sourced on Higgsfield: all prompts and assets are public now. Made with Seedance on Higgsfield. Watch full film below 👇

2744631.5万
Xで開く

すべての有料コース(先着4500名限定で無料) 𝗣𝗮𝗶𝗱 𝗖𝗼𝘂𝗿𝘀𝗲 無料 (PART - 1) 1. 人工知能 2. 機械学習 3. プロンプトエンジニアリング 4. Claude, Chatgpt, Grok 5. データ分析 6. AWS認定 7. データサイエンス 8. ビッグデータ 9. Python 10. 倫理的ハッキング (72時間限定) いいね + RT + コメント「Drive」 DMを送れるように必ずフォローしてね。

1,2042,49930.6万
Xで開く

Hermesエージェントに自分専用のライフOSとセカンドブレインを構築させる方法について、これまで見た中でも最高クラスの解説記事 EPが公開したのは、4つの連携レイヤーを作り出すたった一つのプロンプト: 1. あなたの信頼できる情報源となるMarkdownナレッジベース 2. 今後どのセッションでも場所がわかるようにする永続的なメモリポインタ 3. コンテキストの取得・保存・修正を行うエージェントスキル 4. 習慣・目標・プロジェクト・振り返りをセッション間で同期する、プライベートなライフトラッカー 理解するのはなかなか大変なので、システム全体を図解してみた。保存して、このプロンプトと一緒にエージェントに送ってみて

EP@eptwts

as promised, this prompt will change your life... (send it to your hermes agent & thank me later) ---------------------------------------------- create this system - a generic, durable, private, file-based knowledge base and synchronized life tracker that works across sessions without making the user repeat context. build and test the working system. inspect first and extend any existing knowledge base or tracker instead of creating a competing one. ## core outcome create four connected layers: 1. a portable markdown knowledge base that is the detailed source of truth. 2. a compact persistent-memory pointer that tells future agents where the knowledge base lives and how to retrieve from it. 3. reusable agent skills and scripts for capturing, retrieving, correcting, and maintaining context. 4. a private life operating system for checklists, habits, goal breakdowns, projects, reviews, and read-only life data, with the agent handling planning and contextual writes. all four layers must use the same conventions, dates, privacy rules, and source hierarchy. ## non-negotiable principles - files are the authoritative detailed record. do not use model memory as the database. - persistent memory stores only the knowledge-base path, stable user preferences, and essential retrieval routing. - conversation history is secondary context, not proof of current state. - never invent facts to fill empty fields. - distinguish direct user reports, verified facts, observations, preferences, and hypotheses. - preserve uncertainty and conflicts instead of silently resolving them. - user corrections override summaries while important superseded claims remain traceable. - use exact dates when known and explicitly record approximate date precision when not known. - current summaries answer “what is true now?” while dated records answer “what happened and when?” - markdown is authoritative; indexes, csv, charts, sqlite, and search manifests are rebuildable derivatives. - optimize for fast natural-language capture. do not turn every casual update into a long data-engineering job. - default to one targeted write plus one concise acknowledgement for routine updates. - do not create multiple writable copies of the same fact. - use lexical search, ids, aliases, tags, dates, and one-hop links before considering embeddings. - never store credentials, passwords, cookies, recovery codes, keys, seed phrases, payment authentication, or portal credentials. - do not send private knowledge-base content to third parties without explicit approval. ## phase one: discovery and user rules before creating files: 1. search for an existing knowledge base, notes vault, tracker, or personal database. 2. if one exists, extend it unless the user explicitly requests a replacement. 3. ask no more than four foundation questions unless the answers are already available: - what should the agent call the user, and which timezone should dates use? - what information must never be stored or must be generalized? - which life domain matters most right now? - which existing documents, links, exports, or records should be ingested first? 4. write the answers into the system rules before capturing personal context. done: one knowledge-base location, timezone rule, and privacy/exclusion list are agreed. ## phase two: knowledge-base structure create this portable structure and adapt domains only when useful: ```text life-knowledge-base/ ├── readme.md ├── agent_rules.md ├── 00-index/ │ ├── master_index.md │ ├── open_questions.md │ ├── people_index.md │ ├── projects_index.md │ └── sources_index.md ├── 01-inbox/inbox.md ├── 10-profile/profile.md ├── 20-timeline/ │ ├── timeline.md │ └── events/yyyy/ ├── 30-health/ │ ├── health_summary.md │ ├── conditions/ │ ├── medications/ │ ├── procedures/ │ ├── tests/ │ └── logs/ ├── 40-nutrition/ │ ├── nutrition_summary.md │ └── logs/ ├── 50-projects/ │ ├── projects_summary.md │ └── projects/ ├── 60-finance/finance_summary.md ├── 70-goals/ │ ├── goals_summary.md │ └── goals/ ├── 75-ideas/ │ ├── ideas_index.md │ └── ideas/ ├── 80-interests/interests_summary.md ├── 85-resources/resources.md ├── 90-sources/ │ ├── records/ │ └── files/ └── 95-system/ ├── schema.md ├── changelog.md ├── intake.md ├── backups.md └── templates/ ``` rules for the structure: - keep one concise canonical summary per major domain. - put detailed history in dated event or entity notes, not in giant summaries. - use one file per durable project, goal, condition, medication, decision, or entity when it improves retrieval. - keep one categorized resources file until its size harms retrieval. - preserve original documents in `90-sources/files/` and create a source record in `90-sources/records/`. - put unprocessed material in the inbox, then clear or link it after normalization. - append significant structural or high-value context changes to the changelog. done: every directory has a purpose, all entry points exist, and the master index links every canonical summary. ## phase three: schema and provenance use yaml frontmatter where it improves retrieval. only include applicable fields. ```yaml --- id: evt-yyyymmdd-nnn type: event subtype: health | nutrition | project | personal | finance | goal record_status: current | historical | resolved | uncertain | superseded epistemic: fact | self_report | observation | hypothesis | preference occurred_at: yyyy-mm-dd valid_from: yyyy-mm-dd valid_to: null recorded_at: yyyy-mm-dd created: yyyy-mm-dd updated: yyyy-mm-dd people: [] projects: [] source_ids: [] confidence: high | medium | low verification: verified | user_reported | unverified needs_verification: false supersedes: [] aliases: [] tags: [] --- ``` use stable identifier prefixes: - `evt-` event - `src-` source - `sym-` symptom - `cond-` condition or diagnosis - `med-` medication - `sup-` supplement - `lab-` laboratory result - `per-` person - `org-` organization - `prj-` project - `dec-` decision - `goal-` goal - `idea-` inactive idea - `res-` curated external resource when individual ids are useful record bodies should use this pattern when relevant: ```markdown # human-readable title ## summary ## details ## evidence and provenance ## connections ## uncertainty and follow-up ``` source records include origin, known author or provider, document or event date, date received, path or url, extraction status, and reliability. done: every consequential claim links to a source or is labeled user report, observation, preference, or hypothesis. ## phase four: domain-specific rules ### projects for each active project, capture: - purpose and desired outcome - status and current phase - owner and collaborators - target users - scope and non-goals - architecture or operating model - current priorities - decisions and rationale - constraints, risks, metrics, deadlines, and review triggers - important links, repositories, documents, and source records - next actions without duplicating facts from the canonical project note separate facts, assumptions, forecasts, decisions, and experiments. give time-sensitive metrics an `as_of` date and source. ### resources when the user saves a url: 1. read the original source before summarizing it. 2. store the canonical url and date accessed. 3. record its purpose, useful capabilities or ideas, limits, and future retrieval triggers. 4. preserve installation or invocation commands only when the source documents them. 5. do not treat a saved resource as endorsed or verified truth. 6. reopen the original source before exact quotation, installation, purchase, or consequential use. ### goals capture outcome, baseline, target, rationale, constraints, deadline, measurement, plan, review cadence, progress, and adjustment trigger. distinguish goals, projects, milestones, tasks, and recurring habits. break goals into this hierarchy when useful: ```text life area └── goal ├── success criteria ├── milestones │ └── projects or workstreams │ └── next actions and checklists └── recurring habits or routines ``` every active goal must have: - a clear outcome and definition of done - a small number of measurable success criteria - the next milestone - at least one concrete next action when action is possible - any recurring habits that support it - a review date or cadence - a status such as `not_started`, `active`, `blocked`, `paused`, `completed`, or `abandoned` - a reason for changing, pausing, or abandoning it do not create fake precision. some goals are binary, some are milestone-based, some use a percentage, and some require a qualitative review. ### habits and routines - support daily, weekly, selected-weekday, interval, and minimum-frequency habits. - allow a parent habit to contain individually checkable subtasks. - distinguish a habit from a temporary checklist and from a project task. - record completion only from an explicit user action or reliable evidence. - do not infer completion because a reminder fired or time passed. - permit pause, skip, replacement, and rescheduling without corrupting historical completion. - calculate completion only on scheduled days; unscheduled days are not failures. - avoid worshipping streaks. show consistency, rolling completion, and recovery after missed days. - let habits link to the goals they support, but never claim that habit completion caused the goal outcome. ### general life data support user-defined numeric, boolean, categorical, duration, count, rating, or text metrics, including sleep, mood, energy, symptoms, exercise, focus, learning, social time, finance, body measurements, nutrition, weight, or medication adherence. each metric defines its id, label, type, unit, valid range or options, aggregation, chart, privacy, and whether missing means unknown, zero, not applicable, or incomplete. ### ideas ideas use `not_started`, `exploring`, `promoted`, `parked`, or `rejected`; when work starts, preserve the idea and create a bidirectional project link. ## phase five: capture and retrieval behavior ### capture workflow when the user provides durable context: 1. decide whether it is stable and useful enough to save. 2. search for the authoritative existing note before writing. 3. append to an existing same-day event when appropriate instead of creating fragments. 4. preserve the user’s wording when nuance matters. 5. update the smallest authoritative set of files. 6. update a canonical summary only when the new information changes current understanding. 7. preserve corrections and superseded history. 8. verify with one representative search. 9. acknowledge the capture briefly. for large imports, preserve the original and one source-grounded summary; deeply normalize only when requested. ### retrieval workflow before answering a question about the user: 1. read `agent_rules.md`. 2. search the original knowledge-base files. 3. read the relevant canonical summary. 4. read newer matching events and entity notes. 5. follow source links for consequential claims. 6. distinguish current, historical, resolved, uncertain, and superseded context. 7. state uncertainty and conflicts explicitly. 8. cite note paths, source ids, and dates when accuracy matters. 9. obey the user’s requested scope and exclusions literally. completion criterion: the answer can be traced to current notes and newer records rather than stale memory. ## phase six: persistent memory save one compact durable memory entry that says: - the absolute knowledge-base path - that it is the detailed source of truth - agents must search it and follow `agent_rules.md` before answering personal questions or storing substantial context memory holds only stable routing preferences, never histories, logs, or raw documents. ## phase seven: agent skills create these reusable skills with yaml frontmatter, lowercase hyphenated names, “use when” descriptions, versions, tags, steps, pitfalls, and verification checklists. ### `life-knowledge-base` responsibilities: - search and read personal context - capture durable updates - apply privacy and correction rules - maintain summaries, events, entities, indexes, sources, and changelog - perform evidence-only exports - validate links, ids, dates, csv headers, and retrieval include scripts for deterministic validation and optional lexical indexing. scripts must never print secrets. ### `life-tracker` responsibilities: - read today’s checklist - complete or reopen a habit or nested subtask only from explicit user confirmation - create and maintain recurring habits, schedules, routines, and temporary checklists - connect habits and tasks to the goals or projects they support - log any registered life metric from natural-language reports - add, correct, and remove structured events or measurements - read goal, project, habit, and custom-metric trends - update metric definitions or targets only when the user explicitly requests it - use the user’s timezone for dates natural-language logging rules: - log only what the user says actually occurred or completed, not intentions or reminders. - preserve a useful human-readable description, quantity, unit, category, and source if relevant. - use exact values when supplied. - if missing details materially affect a value, ask one focused question or obtain permission to use an estimate. - mark every agent-derived value as estimated and preserve the assumption. - after a write, report the date, record, estimate status when relevant, and updated progress or aggregate. - after corrections or batches, read the data back from the api. - never guess record ids when deleting; read first. ### `goal-planning` responsibilities: - convert broad ambitions into clear outcomes and definitions of done - identify constraints, baselines, success criteria, milestones, projects, habits, and next actions - keep the plan small enough to execute - schedule reviews and define adjustment triggers - distinguish outcome progress from activity volume - preserve the user’s priorities instead of inventing new goals - update the canonical goal note and tracker hierarchy together without duplicating truth ### `project-context` responsibilities: - create and maintain project notes - break projects into milestones, workstreams, tasks, dependencies, and the next executable action - capture decisions, assumptions, constraints, metrics, links, deadlines, blockers, and review triggers - retrieve current project state before planning or execution - prevent stale summaries by checking newer dated records - archive completed tasks without losing decisions or project history ### `life-review` responsibilities: - run daily, weekly, monthly, and quarterly reviews without duplicating canonical records as cron jobs - summarize completed actions, stalled goals, project blockers, habit consistency, and meaningful metric changes - distinguish observations from explanations - identify a small number of decisions and next actions - update goal or project status only with evidence or user confirmation - avoid guilt-oriented language and avoid turning every metric into a target ### `resource-library` responsibilities: - inspect original urls - create concise retrieval-oriented resource entries - prevent unsupported summaries - search the library before recommending a previously saved tool or source done: each skill predictably changes behavior, uses exact operations where possible, and validates independently. ## phase eight: private life tracker build a minimal private life operating system sharing authenticated server state with agent skills. it guides today’s actions, routines, goals, projects, reviews, and life data; stuff like nutrition and weight is good to track too. use a connected hierarchy rather than unrelated lists: ``` life area → goal → milestone → project or workstream → task └────────→ recurring habit or routine metric definition → dated observations → aggregates and trends ``` keep it flexible: a goal may need only a checklist or several projects and habits. ### tracker information architecture use four primary views: 1. `today` - one focused interactive checklist - previous day, current date, next day - the day’s habits, routine subtasks, due goal actions, and project next actions in one stable vertical list - expandable parent habits with individually checkable subtasks - completed habits in a collapsed section - completion count and progress by day - links showing which goal or project an action supports - concise safety guidance where relevant 2. `goals` - active goals grouped by life area - outcome, definition of done, status, target date when set, and next review - milestones and progress based on explicit success criteria - supporting projects, habits, and next actions - blocked, paused, completed, and abandoned states without deleting history - an agent command or conversation is used to create, restructure, or reprioritize goals; the web app may allow simple checklist completion 3. `projects` - active projects with status, phase, owner, deadline when applicable, blockers, and next action - expandable milestones, workstreams, and task checklists - project health based on explicit rules, not invented percentages - recent decisions and updates linked to the knowledge base - completed and archived projects remain retrievable 4. `data` - read-only dashboard - no contextual log, metric, target, project, goal, or note form inputs - no direct delete or correction controls - the agent is the write interface for contextual life data - compact summaries for goal progress, project delivery, habit consistency, and selected custom metrics - a metric picker generated from the user’s metric registry - one readable chart at a time rather than a wall of graphs - useful chart types chosen by data semantics: line, bar, calendar, count, distribution, or categorical history - range controls such as 7, 30, 90, and 365 days where useful - coverage counts showing how many observations support each average - automatic refresh on focus and a reasonable polling interval so agent writes appear without manual editing the web app handles explicit completion toggles. the agent handles goal creation, project restructuring, metric definitions, logs, corrections, and interpretation using natural language and knowledge-base rules. ### data semantics - every metric declares its own missing-value semantics. - unknown, zero, not applicable, not scheduled, and not completed are different states. - averages use only valid observations unless the metric definition explicitly says otherwise. - intermittent metrics use only recorded measurements; missing days are ignored. - habit completion averages use only scheduled habit days. - goal progress is derived from success criteria or completed milestones, not task volume alone. - project progress is derived from explicit milestones or task scope and must never be fabricated from elapsed time. - paused or intentionally skipped habits are not silently counted as failures. - do not imply that correlations among habits, metrics, goals, health, or projects prove causation. - store timestamps and selected dates separately. - preserve observations, completions, events, and status changes for correction without rewriting history. - optional modules add rules, such as nutrition averages excluding unlogged days and weight trends excluding missing weigh-ins. ### api and security create authenticated endpoints for: - reading dated checklist state - completing or reopening a habit or nested subtask - creating and updating goals, success criteria, milestones, review dates, and goal status - creating and updating projects, workstreams, tasks, dependencies, blockers, and project status - linking habits, tasks, projects, and goals - reading the data dashboard - defining custom metrics and aggregation semantics - adding, correcting, and deleting dated metric observations or life events - changing tracker targets, schedules, and review cadences requirements: - browser sessions use a secure, http-only, same-site cookie. - agent access uses a separate bearer credential stored outside source code. - never expose bearer credentials in browser bundles, urls, logs, docs, or public environment variables. - validate dates, primitive types, identifiers, ranges, unknown fields, and scheduled-habit constraints. - prevent concurrent checklist and agent updates from data loss. - make storage replaceable: local sqlite or json for local use, and an authenticated durable store for deployment. - include backup and recovery instructions. ### design rules - use a restrained neutral interface with readable system typography. - primary text must be readable without zooming. - frequent touch targets must be at least 44 by 44 pixels. - use one centered content column for the checklist. - avoid permanent calendar strips, three-column task boards, repeated summary cards, and simultaneous charts. - show normal sync state as a subtle indicator; show errors only when actionable. - provide explicit loading, empty, offline, retry, and all-complete states. - use accessible names, visible focus, selected states, chart descriptions, and meaningful text alternatives. - avoid decorative copy, excessive borders, tiny uppercase labels, and technical api controls in primary navigation. ## phase nine: tests and validation create and run tests for: - schema normalization and invalid data rejection - canonical-summary and event separation - correction and supersession behavior - privacy exclusions and secret scanning - source and provenance links - representative search retrieval - checklist parent and nested-subtask completion - habit schedules, skips, pauses, and unscheduled-day semantics - goal decomposition, milestone completion, and review dates - project task, dependency, blocker, and status transitions - links among goals, projects, habits, and tasks - concurrent updates without lost state - natural-language agent logging commands - custom metric validation and aggregation rules - missing values following each metric’s declared semantics - intermittent measurements excluding missing days - goal progress not being inferred from elapsed time or unrelated task counts - empty and sparse data - unauthorized api requests - mobile layout without horizontal overflow - minimum touch targets and accessible names - chart metric and range switching - data view containing no form inputs - production build and live authenticated smoke tests also validate: - all required files exist - frontmatter parses - ids are unique - internal links resolve where required - csv headers parse - no secret appears in source or output - the working tree or artifact directory contains every promised deliverable ## final delivery finish by reporting: - the knowledge-base path - the tracker path and live url if deployed - every created skill and its trigger - privacy rules and storage model - tests executed and actual results - backup and recovery method - any assumptions or remaining user decisions provide a short usage guide with examples such as: - “remember that i prefer…” - “save this resource…” - “what do we know about this project?” - “turn this goal into milestones, projects, habits, and next actions.” - “what should i do today to move my active goals forward?” - “i completed the morning routine.” - “mark the research milestone complete and show the next project blocker.” - “create a weekly habit for three focused work sessions.” - “log my energy as 7 out of 10 and two hours of focused work.” - “run my weekly review.” - “show goal progress, project status, habit consistency, and the life metrics i selected.” - “log this meal” or “record my weight” when the user has enabled those optional metrics. claim completion only after the knowledge base, skills, tracker, tests, and verification artifacts exist and run. explain to the user what was just created after creation & how they can use it

1672,07229.1万
Xで開く

すべての有料コース(最初の4500人限定で無料) 𝗣𝗮𝗶𝗱 𝗖𝗼𝘂𝗿𝘀𝗲 無料化(パート2) 1. 人工知能 2. 機械学習 3. プロンプトエンジニアリング 4. Claude, Chatgpt, Grok 5. データ分析 6. AWS認定 7. データサイエンス 8. BIG DATA 9. Python 10. 倫理的ハッキング (72時間限定) いいね + RT + コメントで「Drive」 DM送るのでフォロー必須です

8221,86128.1万
Xで開く

新しいAIテストは「ティール・テスト」にすべきだと思う。入力を「文字通りにではなく、真剣に」受け取るべきタイミングをAIが理解しているか、というテストだ。 どう思う?

1433227.3万