patrick sAIkas

Blog · July 13, 2026 · 5 min read

A development mirror, built from my own emails

I gave Claude six months of my own sent communications, asked it to act as a behavioral assessor, and told it to show me what it saw. The report it returned included a development area no 360 has ever surfaced for me. I couldn't argue with it, because every claim arrived with my own words attached, dated and sourced.

I've spent my career as an I/O psychologist around assessments: personality inventories, structured interviews, simulations, multi-rater feedback. So I want to be careful with what this was and wasn't. It wasn't a validated instrument. It was something closer to a mirror, and it worked well enough that I think anyone who manages people, or wants to grow, should know the method exists.

The idea in one picture: observed behavior, quoted back at you, with the limitations stated.

The setup

The data was six months of my own outbound communication: sent emails from Outlook, my Slack messages, and Zoom meeting transcripts. Deliberately my side only. I wasn't interested in analyzing anyone else, and I'd encourage the same restraint if you try this.

The prompt did most of the work, and its design borrowed directly from how good assessors operate:

  • Observe behavior, then infer patterns. Not "tell me about my personality." Describe what I actually do, repeatedly, in real interactions.
  • Anchor every claim to evidence. Each finding had to cite verbatim quotes, with dates and the source channel. No quote, no finding.
  • Rate your own confidence. Every finding carried a confidence level, with the reasoning behind it.
  • State the data limitations up front. What this corpus can and cannot show.

Those last two matter more than they look. A model that must attach evidence and grade its own certainty produces a very different document than one asked to freestyle an opinion of you. It's the difference between feedback and a horoscope.

One practical note: run this on the most capable model you have access to, with extended thinking turned on. The whole value is in how much data it can hold and cross-reference, and how much thinking about that data the model can do.

The finding that stung

The report counted the number of times I described AI development work as "trivial" in messages to colleagues. It quoted the instances back to me, with dates. Then it made the case: tasks I've done dozens of times aren't trivial to someone doing them for the first time, and calling them trivial sets those people up to feel stupid when the "trivial" thing takes them a full day.

I knew the curse of knowledge as a concept. I did not know I was performing it in writing, on a regular schedule, with a countable frequency. That's the part self-report can never give you: I would have rated myself well on "sets realistic expectations for others." The data disagreed, and the data had receipts.

Why this is different

Our field has spent decades getting better at two kinds of signal: what people say about themselves, and how people behave in simulations built to provoke behavior. Multi-rater feedback added other people's perceptions, which are valuable and also filtered through memory, relationships, and politics.

This is a third kind of signal: actual behavior, at scale, passively collected, fully auditable. Nobody had to remember anything. Nobody had to soften anything for the relationship. And when I doubted a finding, I could click through to the primary evidence, because the primary evidence was my own sent folder.

The limits, stated plainly

The report itself said this, and I'll repeat it: this is not a validated instrument. There are no norms, so "high" or "low" means nothing except relative to my own baseline. There are no rater perspectives, so anything that lives in how people experience me in a room is invisible. The corpus is written and transcribed communication, which oversamples certain behaviors and misses others entirely. And a language model's inferences can be confidently wrong, which is exactly why the prompt forces evidence and confidence ratings you can check.

So I treat every finding as a hypothesis, not a verdict. The useful ones are the hypotheses that sting a little and survive your own audit of the quotes.

If you want to try it

The recipe, in short: gather a body of your own sent communications, the more months the better. The gathering is easier than it sounds, because you don't need to export anything. Link a model up with read-only access to wherever your communication lives (Outlook, Gmail, Slack, Teams) and tell it to look only at your sent messages. Use the strongest model available to you with maximum thinking enabled. Ask it to act as an expert behavioral assessor. Require verbatim quotes with dates and sources for every finding, require a confidence rating on each, and require a stated list of what the data cannot support. Then read it the way you'd want a coachee to read a 360: look for the pattern you'd have argued against yesterday.

Honest question, the same one I asked when I first posted about this: would you want to read what your communications data says about you? I've found most people's answer is a fascinated, slightly nervous yes. That reaction is usually a sign it's worth doing.