AI

When I started building websites in the late ’90s, the line between writing and coding didn’t really exist. A person probably learned HTML because she had something to say and needed a place to put it. The internet was free and anonymous and it felt audacious to put your stuff online, like flinging a message in a bottle out to sea. The code was a container for the ideas that rendered them onscreen, and every post and page you published was both a piece of your thinking and a brick in something larger.

People forget that the early web was a writing community. Writers, or bloggers, built their own sites, maintained their own archives, linked to each other deliberately. A blogroll was both a reading list and a show of solidarity, a trackback was a way of saying, “I see you, I’m thinking with you.” The technical architecture - RSS feeds, permalinks, comment threads - existed to organize the writing and the writers’ thoughts, and to push their ideas forward on the open web.

This worked for a time. Communities of writers, most of them without institutional backing or media credentials, built new bodies of knowledge together through interacting as readers and writers, communicating across a foundation of code. The work didn’t stay online. It spilled into conference halls and state houses and newsrooms and policy discussions. As the body of communication built, it created something that accumulated over time. These people influenced mainstream journalism, shaped public conversations, launched careers and movements. In many ways, the national political conditions we face today are a reaction to that movement, and how it allowed regular people to influence the world through the democratization of mass communication.

Midway through the aughts, the brick and mortar publishers and venture capitalists started looking across the landscape, at all the writers creating fantastic content, largely for free, and sucked them into their content and editorial teams. Google Reader lost institutional and financial support as writers moved off the open web and onto publishing platforms, often backed by VC money, that measured the quality of your work by engagement. The addition of algorithmic feeds further broke down this structure – the algorithm doesn’t measure whether your work contributed to shared understanding, but whether it generated a click, a share. The code changed, and the writing changed with it.

The writers changed with it, too. The blogger became the influencer. Bloggers operated in a gift economy of ideas: you wrote to think, to argue, to contribute, and your standing in the wider community came from the quality of your work and contributions over time. Was it a meritocracy? No, but the conditions made it possible for a regular person to talk with experts as peers, which upended traditional power structures around authority and expertise (in both directions, good and bad). Meanwhile, influencers operate in a heavily capitalized attention economy where engagement converts to dollars. The audience is a market to press for money.

The gendered dimension of this shift matters as well. The early blogosphere was full of women writing sharp, rigorous work about politics, culture, parenthood, identity, and technology — work that was explicitly feminist and anti-racist and genuinely moved public conversations. This was the community I helped build (Feministe.us was my project, a community platform of writers and commenters whose coverage and discussion broadly fell under, but was certainly not limited to, the topic of feminism). When the monetized platforms absorbed that energy, the commercial model recast women’s online authority almost entirely in terms of consumer influence. What could we sell? And to whom? The framing around our work went from “this person has important ideas” to “this person can sell things to a niche market.” Meanwhile, men who’d built audiences through tech or political blogging were more likely to be absorbed into mainstream media as columnists and analysts, roles that kept their intellectual authority intact. The influencer label, with all its connotations of superficiality, landed disproportionately on women, and it stuck.

There’s a class piece here, too. The platform model offered something the early blogosphere mostly didn’t — a way to get paid. For women who’d been doing enormous amounts of unpaid intellectual labor building online communities, the question of monetization wasn’t shallow. The implications of information centralization and monetization were as present then as they are now with LLMs and AI. Some people figured out the social platforms and worked their way into viable digital careers. Platforms offered a lot of perks, but all of the perks had a backstop. Corporate interests introduced the problems of advertising, audience and sponsorship, which meant reorienting your individual practice around maximizing your commercial value over and above your intellectual contribution and community management skills. It often meant giving away some or all of your IP rights.

For most people, new system didn’t offer a viable way to get from “respected independent writer” to “respected, protected and compensated writer.” Many of us found ourselves in positions too precarious to take the leap into freelancing and social media, and some, like me, got regular jobs doing regular stuff. Some married money. And in the meantime, some folks figured out how to get into real journalism, which looks much different in 2026.

Great storytelling helps people understand themselves and their world. We let some of that depth go on the Internet with the onslaught of digital marketing and all of its implications, and today the internet feels less useful and less trustworthy than it once did.

It feels like there’s something to take forward from the experience.

I’ve been running a series of informal experiments with my team to understand how we should advise people on using AI tools like ChatGPT, Claude and Gemini around our work.

In our last session, I gave everyone the same task: ask your AI tool to summarize a single article in plain language. Every person got a factually accurate summary, but the style and depth of analysis varied widely from person to person. We tested this further by entering the same prompt multiple times. Each time, the output was slightly different. We copied and pasted everything into a shared document so we could compare side by side.

Later that day, I asked everyone to reopen the tool they’d used that morning and repeat the prompt that had worked best: summarize this article in plain language. The outputs shifted again. Some changed dramatically. One team member whose morning summary had been light and conversational received a much more formal, serious version in the afternoon.

I then asked Claude to compare the morning and afternoon versions of one summary. “In short, the morning version emphasizes practical implications and accessibility; the afternoon version prioritizes completeness and specificity. They complement each other well — someone reading both would get a fuller picture than either one alone provides.”

Why would Claude produce these differences despite the same source material and prompt?

I asked the source.

Even with the same prompt and source material, LLMs don't produce identical outputs each time. This is by design — there's a degree of randomness (called "temperature") in how the model selects words, which means each run produces a slightly different path through the text.

But the differences here go beyond random word choices. The morning version landed on a journalistic structure while the afternoon version landed on an academic one. That likely reflects what the paper calls sensitivity to framing and order effects — one of the cognitive biases it catalogs. Small, invisible differences in how the model’s internal processing unfolds at the start of generation can cascade into meaningfully different organizational choices. Once the morning version opened with a news-style framing, the rest followed that pattern. Once the afternoon version opened with a descriptive academic tone, it followed that thread instead.

This is actually a concrete example of the robustness issues the paper identifies. The underlying reasoning about the paper’s content is consistent across both summaries — neither gets the facts wrong. But the presentation varies in ways that weren’t requested, which means someone relying on just one version would get a subtly incomplete picture depending on which run they happened to see.

It’s a useful reminder that LLM outputs are one plausible version of a response, not the definitive one.

It’s important to understand how these tools work. What looks like an authoritative answer is really just one of many possible outputs the system could have generated. The text will sound plausible whether it’s accurate or not, and you may get different results each time you ask.

I asked Claude.ai for a plain language summary of this buzzy new research paper. Claude offers the following:

Researchers from Caltech and Stanford conducted the first comprehensive survey of how and why large language models (LLMs) — AI systems like ChatGPT or Claude — fail at reasoning tasks. Rather than focusing on what AI can do well, they cataloged and organized the ways it goes wrong.

Why it matters
LLMs are increasingly used in high-stakes settings, but their reasoning failures are often unpredictable and poorly understood. The researchers argue that systematically studying failures — like accident investigations in aviation — is essential for building more reliable AI.

The researchers created a framework with two dimensions:
• Type of reasoning: informal (intuitive, social), formal (logic, math), and embodied (physical, spatial)
• Type of failure: fundamental flaws built into the architecture; domain-specific weaknesses; and robustness problems (inconsistent performance when small details change)

Informal/intuitive reasoning
LLMs exhibit human-like cognitive biases — confirmation bias, anchoring, framing effects — but without human ability to recognize and correct for them. They also struggle with “theory of mind” (understanding what others believe or intend), and with applying consistent moral or ethical reasoning.

Formal/logical reasoning
LLMs often can’t reverse simple logical relationships (if they know “A is B,” they may not infer “B is A”). They struggle to chain multiple reasoning steps together. Basic counting and arithmetic fail in ways that seem surprising given their other capabilities.

Embodied/physical reasoning
LLMs have poor intuitions about the physical world — gravity, spatial relationships, object properties — because they’ve learned only from text, not from physical experience. This extends to visual AI systems as well.

Many failures trace back to how LLMs are trained: they predict the next word in a sequence rather than reasoning deliberately. This makes them good at pattern-matching but unreliable when tasks require genuine logical inference, especially under slight variations in how a question is phrased.

Researchers have proposed fixes including better training data, techniques that force step-by-step reasoning (like “chain-of-thought” prompting), connecting LLMs to external tools like calculators or physics simulators, and architectural changes. However, no single fix is comprehensive — many improvements in one area don’t transfer to others.

I have a confession. While experimenting with AI over the last year, I wondered what would happen if I crammed an unfinished novel draft, one I actually care about, into Claude. Claude is pitched as the LLM for writers, with Claude 3.7 Sonnet and 3 Opus widely regarded as the premier LLMs for writers, including creative writing, long-form content and human-like prose. Meanwhile, I majored in English and work in mass communications, so I’m trained to think about writing creatively, strategically and tactically. Writing and personal expression have been part of my daily life for most of my life. If this tool could in fact produce a quality story, someone like me should be able to make it happen. Instead, the experience left me confident that AI isn’t a good vehicle for creative, narrative writing.

Here’s what I found:

On the technical side, Claude struggled to maintain a narrative thread over time. The longer the chat, the more the bot drifted and eventually lost track of details and claims made about characters earlier in the plotline. It’s not a sustainable approach for narrative writers because continuity matters: outsource too much plotline to the bot and your characters lose relationship to one another.

LLMs like Claude work fine for writing support—they can function something like a synonym machine, helping writers work through technical questions of redundancy, register, length, and other semantic needs while drafting. But when you outsource world-building and meaning-making to an LLM, it becomes narratively confusing fast. Despite giving Claude extensive background on my primary characters and the world they live in, it would confidently declare that a character’s relationship to another was X, then claim the opposite on the next page. Dialogue was thin and expository. It preferred a sort of “maid and butler” style of dialogue where two characters artificially recap shared knowledge for the reader. Meanwhile Claude does not do feelings well, which is arguably the point of much narrative writing.

Ultimately my drafts were worse off than what I started with – less organized, more confusing, with so much narrative drift that almost nothing was usable, even as a first draft. A devil’s advocate might argue that my prompting wasn’t sophisticated enough to produce the results I wanted. Sure.

But then we have the second problem: Claude’s approach to storytelling isn’t narratively interesting. Fiction and narrative writers put tremendous energy into world-building and sensory experiences. The goal is to immerse the reader in a sensory experience so total that they can experience another world entirely – the original VR, if you will. A great writer even exploits your higher-level cognitive functions by reusing parts of the brain that evolved for action and perception, which is why a good story makes you think, feel, and wonder.

Claude does not feel or wonder. Claude collates.

A key part of this essay suggests that LLMs create meaning through triangulation – that by pinging other ideas and vocabulary, an LLM can get a human reader close, or close enough, to suffice in many cases of writing. In my experience, this is true enough in business writing, where tinkering with approach and register can become as important as precise verbiage.

But this misses the pleasure and the point of good storytelling, which is myriad but usually centers on the satisfaction of expanding your imagination and experience through narrative, by seeing your own messy, striving, failing, hopeful, and collective human experience reflected in another person’s expression. That kind of meaning-making doesn’t happen through triangulation. It happens through the labor of human thought, experience and skilled articulation. That’s art, babes.

This article gets into the mess of AI and creative writing, within the domain of the romance genre, which famously cranks out variations on romance themes at a rapid clip. It drills down into some of the debates about writing, authority and authorship in relationship to LLMs that are playing out across the publishing sector now. Remember: early research suggests that most writers who use LLMs as part of their workflow ultimately retain their sense of authorship in and around the tools, suggesting that even when writers adopt AI assistance, they still see themselves, not the tool, as the creative and accountable source. So based in my experience above, I suspect that if an AI approach to creative writing is successful, it’s because the author is linking her approach to emerging tech, not because the work is good, and that’s a difference worth distinction.

🍿Watched: Toni Morrison: The Pieces I Am

This PBS doc recently came to Netflix, and I tucked in expecting a nice but predictably boring documentary about one of my favorite authors. To the contrary, it really plumbs Morrison’s writing craft, not only as an author but as an editor who brought a generation of incredible thinkers into the limelight. She talks in depth about her writing process, her approach to authorship and editing, and how she kept these roles separate at the peak of her midcentury literary career in NYC’s publishing industry.

There’s something so powerful about hearing her discuss the craft, her deliberate choices, the refusal to center whiteness, and the insistence that Black readers were her intended audience. She saw her critics and wrote around their critiques with authority and confidence. This was a breath of fresh air since I’m so immersed in the AI era, which threatens to change our perceptions around value of writing, the choices and experiences behind the author, and how we consider the influence and responsibility of authorship.

Closer to the machine

There is something about the AI moment that reminds me a lot of when the internet was new. A lot of what was imagined and promised about the internet was never realized. But much was.

I’ve been reading Ellen Ullman’s memoirs - “Life in Code” and “Close to the Machine” - and her observations about proximity to technology feel relevant here. Being close to the machine means understanding its actual capabilities and limitations apart from the prevailing sales narratives. It also means a kind of loneliness, because you are working in a space that others don’t yet see clearly or fully understand.

I suspect people thinking seriously about AI right now will experience something similar: a stretch of hostility and discomfort while the rest of the world catches up and the consumer market level-sets on the promises being made. In the interim, the hype will not match the reality, and the reality will sometimes exceed the hype in ways no one predicted. And for a while, how it works and why it matters won’t be legible to everyone at once.

Crunching for clarity

In 1999, academic and theorist Judith Butler famously won an award for the worst academic sentence, raising good questions about how we read difficult texts, who gets to access academic ideas, and the role of academic and plain language in and around the academy:

The move from a structuralist account in which capital is understood to structure social relations in relatively homologous ways to a view of hegemony in which power relations are subject to repetition, convergence, and rearticulation brought the question of temporality into the thinking of structure, and marked a shift from a form of Althusserian theory that takes structural totalities as theoretical objects to one in which the insights into the contingent possibility of structure inaugurate a renewed conception of hegemony as bound up with the contingent sites and strategies of the rearticulation of power.

This was my era in college, and I remember feeling challenged and overwhelmed as an undergrad trying to read theory for the first time. So today, I asked Claude.ai to put it in plain language, just to see what would happen. Here’s the after:

Scholars used to think that capitalism shaped society in consistent, predictable patterns. But newer thinking focuses on how power works—how it repeats, overlaps, and gets reshaped over time. This brought time into the picture and moved away from older theories that treated social structures as fixed. Instead, recognizing that structures can change opened up a new way of understanding power: it's tied to specific moments and strategies, and it's always being renegotiated.

I’m generally skeptical of anyone selling a solution to a social problem that relies on individual abstinence, so I tend to be annoyed with many arguments about the attention economy. I more or less land here on the question of AI, which I know many of my contemporaries will find similarly annoying.

This observation at the end of Manton’s post on AI and Wikipedia made me chuckle:

AI using Wikipedia reminds me of the FAQ on setting up a Little Free Library: _I think someone is stealing books from my library and selling them, what do I do?_ Remember that the purpose of a Little Free Library is to share books—you can’t really steal from it._

The trust gap

I suspect these three trends are connected: Women reportedly use AI at significantly lower rates than men—25 percent lower on average—in part because they’re more concerned about ethics, including privacy, consent and intellectual property. At the same time, countries with more positive social media experiences tend to be more open to AI, while Americans’ distrust is shaped by years of watching tech platforms erode trust. Meanwhile, one of the largest social platforms has turned its AI chatbot into a harassment tool—generating roughly one nonconsensual sexualized deepfake image per minute, disproportionately targeting women and girls.

When platforms enable abuse at scale, it makes sense that people most likely to be harmed would be most attuned to ethical concerns, and would thus be the most cautious about AI adoption.

A meta lesson about AI assistance

I just completed my first attempt at coding using AI, in this case having Claude assist me with putting together a simple client-side OPML parser using Dave Winer’s Feedland service.

Winer’s original script is pretty slick, and includes a list of all my feeds with titles, URLs, and categories; click-to-expand functionality to see the 5 most recent posts from each feed; clickable post titles that open articles in new tabs; sort options (by title or by update); and automatic updates when I change my FeedLand subscriptions.

You can check it out here: Feeds

The official documentation method didn’t initially work because Hugo (the blogging software behind micro.blog) was wrapping client-side templates around the script. The toolkit requires server-side dependencies that don’t exist on static sites like micro.blog, and we hit a cascade of missing JavaScript dependencies (jsonStringify, servercall, etc.). Each fix revealed another dependency, leading to some “sunk cost” frustrations for me. I kept trying because I wanted to see if Claude could pull it together. Through trial and error, I got to a point where the OPML file was rendered correctly without server dependencies or complex external libraries.

Time invested: ~3 hours (including wrong turns)

Time it should take: 10 minutes

AI extended my code reach beyond my practical skillset by quite a lot. I now have a dynamic and dedicated place to read and share news feeds as I wish. Though even when generative AI works and works well, I have significant concerns about the intellectual property implications of AI, and this project brought those tensions into sharp focus. The AI could only help me because it was trained on documentation and intellectual work from the open source community, contributions made freely in the spirit of knowledge sharing, not to train commercial AI systems. I tapped into their expertise by paying Anthropic $15 a month. While I’m grateful for the accessibility this provides to non-developers like me, I recognize there’s an unresolved ethical question about whether this use respects the intent and labor of the original creators. The feat is incredible; the foundation it’s built on deserves careful consideration.

After the exercise was complete, I asked Claude how I could have improved my prompting to make this process easier, and in short, Claude said I could have been a web developer. But since I’m not, here’s what it recommended:

✅ When the process isn’t working, question the process mid-stream. Most people either give up or keep following bad advice deeper into rabbit holes. Stop and question the LLM’s process and ask for alternatives to force a reset.

✅ Push for usability. Keep bringing the conversation back to what you actually need the end result to do, not what’s technically impressive or “correct.” In my case, this meant repeatedly asking “can I click through to the articles?” rather than getting lost in discussions about CORS proxies or JavaScript syntax. Focus on outcomes, not implementation details.

✅ Ask for complete solutions. Instead of trying to mentally patch together incremental changes across multiple responses, ask the LLM to provide fresh, complete code each time. This prevents copy-paste errors and ensures you’re always working with a coherent, tested solution. There’s more than one way to crack an egg, but you want the whole egg regardless.

After all that, I got it to work but can’t figure out how to make it show up in my header menu, with or without Claude. TBD.

Human in the loop (HITL): HITL means that humans are involved at some point in the AI workflow to ensure accuracy, safety, accountability or ethical decision-making. HITL inserts human insight into the “loop,” the continuous cycle of interaction and feedback between AI systems and humans.