Within this report, the LLM reflected my concern about my old blog appearing to be the center of these old networks. In fact, this is an effect of me pouring into the model. It says:
The compiler appears on this page more often than her place in the blogosphere warrants. That emphasis is an effect of how large language models work: the models that compiled the page (Gemini and Claude) build their output from the material in front of them, and the compiler’s recollections were the fullest source they had, so her name attached to claims as their source and her own path filled more of the page than the record supports. This version names her less and marks firsthand claims as “firsthand.”
I asked Claude why LLMs work this way:
Context outweighs training data. A model has two sources: what it absorbed in training and what sits in the prompt. For feminist blogs of the 2000s, the training data is thin and scattered. Many of those blogs are dead links, partial Wayback captures, or sites that only show up as names in other people’s blogrolls. Your recollections, by contrast, were dense, specific, and right there in the context window. Generation is conditioned on that context, so the model draws the words, names, and connections it gives the highest probability from the prompt far more than from faint training memories. The fullest source sets the terms for everything else.
Source and subject merge. When you write “I remember that X linked to Y,” the text contains two facts: something happened between X and Y, and you are the one reporting it. [LLMs are semantically responsive to triangulation of meaning ~LB] Models don’t track provenance well, so the narrator’s name tends to drift into the content. “Compiler recalled” turns into “Compiler was part of,” and the person reporting a network becomes a node in that network. Your note’s phrase “her name attached to claims as their source” describes this accurately. It is a familiar failure in summarization, and it gets worse as the documents get longer.
Detail density reads as importance. Where the record is thin, the model fills gaps by extending whichever thread has the most material, and that thread was your path. A blog you described in four sentences gets more space, more connections, and more apparent centrality than one you mentioned in passing. The page ends up showing how much the compiler knew about each blog. That is a different quantity from each blog’s place in the network, and nothing in the process separates the two.
Training data favors the same names. Feministe is one of the better-documented blogs from that period. It was archived, cited in academic work, and linked from other surviving blogs. So the model’s background knowledge already over-represents the blogs that left records, and your recollections reinforced that bias instead of offsetting it.
Tuning rewards engagement with user material. Preference training rewards responses that use what the user supplied. A model that set aside half your recollections as over-represented would probably have been rated lower during training, so it leans toward incorporating whatever it was given.
Two models compound the effect. Gemini’s output became Claude’s input. The second model received a document that was already weighted toward you and treated that weighting as part of the record. Each pass inherits the previous pass’s emphasis and adds no correction, because neither model has an outside reference point.
The compiled page presented one person’s situated knowledge as a neutral overview, which is close to Haraway’s “god trick.” Labeling claims “firsthand” puts the situation back on the page. For the notes you plan to write, a useful test is whether that label changes how the model weights claims when it compiles the next version, or only how readers weight them. I expect it changes only readers’ weighting unless the prompt also tells the model to give unsourced firsthand claims less space than claims with external citations.
If I were trying to explain this to someone back home, I’d say the business model inverts and empowers compute across the domain of “word processing,” which includes semantic code. And because of the dynamics around the single-user interface, it lends more data authority to those outputs than is warranted. At the same time, there is obviously demonstrable value here, but it remains loose and weirdly ungovernable - like the ghosts in the machine are perhaps us after all.
Access to tech: Enterprise level. Claude Pro plus added tokens (~$40 out of pocket after I ran through Claude Pro’s compute).
Process: Gemini pulled a roster based on the Roberts article in Jezebel. I strongly suspect my call for a “roster” shaped the inputs and outputs, making this cohort of writers look like a crew of fantasy football players or Marvel characters. All stats.
I asked Claude to analyze the data from my perspective, which includes people and details that aren’t part of the formal record, allowing for compare and contrast. To get here, I had to publish and republish the prompt almost forty times to better anonymize the data while protecting the writers included.
The “thickness” heuristic applied is totally made up - Claude took my thoughts about Donna Haraway and McMillan Cottom’s work and wove them through my thoughts about information architecture that I first learned from Dr. B.
Claude scanned the final draft for legal and privacy concerns. Everything here is publicly available online today, because it is currently live or lives somewhere as a ghost in the machine, whether via Wikipedia or the Wayback Machine or otherwise.
Blog note: Because of how my blog is set up to automagically find Wayback links for old URLs, if I publish my reference list, it would open a social and informational can of worms. So the (extensive) reference list will continue to live on the Claude report rather than being indexed into my link library.
I gave Claude permission to look across public data, crack windows and open doors, and to annotate and proceed as it worked, rather than keeping the decisions with me, then shaped the final strawman into what’s here. This took about half a day to pull together - yammering at the data within the Claude interface until it looked reasonable. Imagine having resources and a data scientist apply themselves here, with long reach and memory, with implications for trust, safety, consent and memory in every direction.
In the meantime Claude shut me down for the week (too much compute). I had to wait until this morning to ask Claude how much energy it took to run this report as I did, as a rookie and everyday practitioner.
Does this feel slimy? I concede it may. But it demonstrates the value of personal writing and thoughtful design in digital spaces despite its many gaps and flaws.
Enjoy and tinker. Feedback is welcome. I’ll have some notes on my approach and why many of these findings remain heavily qualified in the coming weeks.
Konstantinos Komaitis has started a ten-part series at Techdirt called The Metric Is Not the Mission. It traces how Big Tech went from building the open internet to reshaping it around its own metrics and incentives. Part I argues that the recurring fights over content moderation, privacy, recommendation algorithms, and AI are symptoms of one long shift.
Search, social networks and recommendation systems solved real problems of navigation and abundance. Eventually each solution becomes the environment itself, and continues to be optimized for incentives that benefit the companies behind tech solutions. His phrase for how we’ve been reading the news about tech is “paying attention to the weather while overlooking the climate.”
The historical framing will only pay off if later installments name specific mechanisms and the people who made the decisions, and I’m reading the rest to see whether they do. New parts come out twice a week, with a full PDF at the end.
While I still go back to Haraway’s “Situated Knowledges” for the argument that every view comes from somewhere, these days, the writer I read most closely on technology is Dr. Tressie McMillan Cottom.
Haraway asks where a knowledge claim is standing, while Dr. Tressie asks who pays for it. Lower Ed traced how for-profit colleges sold credentials to audiences the rest of higher education had shut out. In a TikTok mini-lecture on AI, politics, and inequality, she carried that analysis forward using Daniel Greene’s “access doctrine,” which holds that the answer to economic inequality is more skills rather than more support, and which universities facing their own precarity have leaned into. Audrey Watters and Ed & Class both wrote it up.
Her question fits my work life better right now while AI begins to show up in procurement, policy and training decks, and her ongoing attention to pedagogy keeps the conversation tied to what happens between a teacher and a student. I read Haraway for the epistemology and Dr. Tressie for the current evidence of how it plays out in real classrooms.
Noted: Because I’m using Gemini to find and complete resources related to the earliest blog years, I’m citing a Clancy throughout my project, which the LLM keeps pulling back to center events and articles related to the Clancy trial.
Neither company publishes a timeline, and most of the numbers floating around come from SEO consultants, so treat this as rough. The short answer is that Gemini usually gets there faster because it searches Google’s index, while Claude appears to search Brave’s, which fills in more slowly.
Gemini grounds its answers in Google Search, so its lag is roughly Google’s indexing lag. That can mean a few days for a page on an established site and two to six weeks for everything else. If your robots.txt blocks Google-Extended, Gemini won’t use your pages for grounding at all.
If you paste a URL directly into either chat, the model reads the page right away without waiting on any index. To check whether you’re findable, search your domain on search.brave.com and in Google Search Console. Blocking ClaudeBot, the training crawler, doesn’t stop Claude from citing you; Claude-SearchBot is the one that governs search.
Fellow SEO/GEO nerds are citing an average of six days to show up in Google’s systems, though it is sometimes taking as long as thirty days on smaller, less-visited sites.
Machines and machine learning change the ways we think about humans and humanity. In cyberspace, we create “digital doubles,” or data doubles, or digital human twins (DHTs), and now these digital doubles move through assemblages of AI systems. New politics are emerging from this tangle.
(I am sad to hear that Soraya passed away last month. During her research for this article, I enjoyed discussing the technical details that created these spaces and how our collective work facilitated everything that came afterward, or didn’t. I know Soraya was hanging a lot of her career on this article and felt it was poorly handled at the time, speculating it was because the voices most implicated by the research now manage the levers of publishing.)
Knowing Google has the juice for search, including all the missing bits deprecated over time, I asked it what technical and social lessons were learned, assuming Gemini could treat that work like a portfolio and provide a business analysis with the full benefits of hindsight.
Per Gemini, links added by me:
Looking back at the mid-2000s, the feminist blogosphere achieved extraordinary cultural breakthroughs, but it also operated with distinct technical and social failings and prejudices. By treating cyberspace as a purely ethereal or detached realm, bloggers unintentionally reproduced offline power dynamics, a phenomenon known as the “digital twinning” of cyberspace.
Viewed through a critical, retrospective lens, four major areas show where different choices around technology, institutional strategy, and social solidarity could have altered digital feminism’s trajectory.
1. Technical: platforms are not neutral
The early blogosphere largely treated technical tools (like WordPress, Blogspot, Google Search, and ad networks) as neutral pipes rather than profit-driven entities with embedded incentives. Bloggers built incredible social capital on free or cheap third-party services without owning the underlying code or digital infrastructure. When Web 2.0 shifted toward algorithmic feeds and centralized social media (Twitter, Facebook), bloggers willingly migrated their engagement to those platforms to chase larger audiences.
Protocol-Level Investment: Instead of relying on proprietary platforms, the cohort could have invested in building and maintaining independent digital commons, including federated software, cooperative hosting collectives and shared digital archives owned directly by the movement.
Data Sovereignty: Early writers rarely anticipated how their essays, comment sections and personal narratives would be scraped, monetized, and eventually used as training data for commercial algorithms and search engines. Establishing strict open-source licensing (such as Copyleft or peer-production licenses) and sovereign archiving practices could have protected intellectual labor from platform extraction.
Many of the highest-profile blogs were anchored by white, college-educated, urban writers who inadvertently treated their lived experiences as the “default” feminist lens. Marginalized writers, particularly Black, Indigenous, AAPI, disabled, and trans writers, frequently had to demand inclusion, fight for citation, or build separate spaces (Racialicious, Crunk Feminist Collective, Reappropriate, FWD/Disabled Feminists) to address the dominant blogs’ weak spots.
Resource and Power Redistribution: Rather than treating traffic and ad revenue as individual metric wins, major sites could have operated as resource-redistribution networks—directly funding, platforming, and sharing domain authority with writers from marginalized backgrounds from the outset. This is what Feministe.us aimed to do with our rotating lineup of guests, with mixed results.
Rethinking “public” vulnerability: The blogosphere heavily rewarded raw, highly personal exposure (“confessional blogging”). However, this cost was asymmetric: white, middle-class bloggers often parlayed personal exposure into book deals, while marginalized bloggers faced severe offline economic risk, doxxing and targeted harassment for sharing the same level of personal detail. Many of us also needed to earn a paycheck with benefits. Depending on your demographics, angles and timing, writing opportunities were limited, and only a small fraction of the talent generated in the collective saw durable career impacts. A more intentional approach would have prioritized pseudonyms, pen names, shared authorship, and safety-first rhetoric.
3. The siren song of individual upward mobility
The most consequential structural shift was how the blogosphere interacted with traditional media and corporate capital. As more corporate outlets (Jezebel, HuffPost, The Guardian) realized feminist commentary drove clicks, they offered individual bloggers lucrative staff jobs, book deals and column spaces. This transformed a collective political movement into a pipeline for individual career advancement. As top talent moved into legacy media, independent blogs lost the labor needed to sustain them.
Building cooperative media outlets: Instead of individual writers accepting staff roles at corporate outlets, the cohort could have formed reader-owned media cooperatives (similar to worker-owned publishing models today). This would have allowed writers to earn a sustainable living while keeping their intellectual output in a community-controlled space.
Resisting algorithmic speed: Legacy media required bloggers to output somewhere between 3-10 quick, hot-take posts per day to satisfy ad-driven impression quotas. Resisting this pressure in favor of slower, deeper, “thick” theory would have preserved the intellectual rigor of the spaces instead of reducing feminism to consumable digital commentary.
4. Politics: Hesitancy to be the “fly in the ointment”
Beyond technical tools and corporate pipelines, there was a profound rhetorical hesitation within many of the era’s flagship blogs to comfortably occupy the role of the “fly in the ointment,” the deliberate, disruptive presence that insists on disrupting false consensus to center genuinely aggrieved interests.
Because major blogs faced constant, bad-faith attacks from right-wing media and anti-feminists, an implicit pressure emerged to project a unified “united front.” This often created an environment where internal dissent, particularly from disabled, working-class, trans, or non-white women pointing out the movement’s ongoing failures, was framed as unhelpful infighting or a distraction from the broader political battle. The push for a polished, accessible and broadly appealing “Feminism 101” frequently smoothed over the jagged edges of critique.
Sanctifying disruption as methodology: The blogosphere could have explicitly framed the “fly in the ointment” not as a threat to solidarity, but as the primary engine of accountability. Rather than managing or policing internal conflict to maintain a clean editorial brand, blogs could have deliberately designed their platforms to hold open space for radical and uncomfortable grievances within the plurality of feminist thought.
Protecting uncommodifiable voices: Corporate platforms favor palatable, easily shareable commentary by design. By shying away from deeply uncomfortable, non-mainstream grievances, the blogosphere left a vacuum filled with endless explainers. A more resilient model would have actively protected and amplified the most uncompromising, “unmarketable” voices, including those whose positions could not be sanitized for an ad sponsor or a cable news segment.
By failing to fully institutionalize the “fly in the ointment” as an essential, protected role within digital spaces, the early blogosphere left itself vulnerable to the corporate absorption that eventually dismantled it. When mainstream media stepped in, it easily plucked away the polished, palatable commentary while abandoning the uncomfortable, dissenting labor that keeps a political movement expansive and honest.
Noted: I’ve been playing with Copilot, Claude and Gemini in tandem, and Copilot and Gemini are suddenly, curiously outperforming Claude by a lot.
Eli Lilly redux: Thinking about the democratization of information and the removal of gatekeepers alongside this story I saw first at Boys’ Club. This Harvard PhD is vibe-manufacturing schizophrenia drugs in his garage. Are we vibe-coding medicine?
Manton on AGI: “As we reach AGI… I’d like to set some ground rules for myself — beliefs that won’t change as technology changes. AI will be smarter than us, but consciousness and ‘a soul’ can’t be created from computation.”
A composition teacher friend shared this paper on social media: Lester Faigley’s “Literacy after the Revolution”, the essay version of his 1996 CCCC Chair’s address. In it, the author argues that the economic impacts of the digital revolution had begun to undo an older commitment, formed in the Civil Rights era, to teaching literacy as a path toward equality. He further argues that writing instruction was being reorganized around tools owned by a few firms (then: Netscape, Microsoft) at a moment when wealth was concentrating upward. Faigley left us with the question of whether educators can hold onto literacy-for-equality while the tides run against it.
Thirty years later, the worry has a new face: AI will do young people’s writing for them and their thinking with it. Ultimately, Faigley believed the need for the skills that composition teaches will keep growing, not despite, but because of our need to convey information in and around that technology and the humanity it serves in a complex society. Does that suspicion hold water today?
To read: this report from Common Sense Media, “Talk, Trust and Trade-Offs: How and Why Teens Use AI Companions.” About a third of teens have turned to an AI for serious conversations about their own mental health and relationships, showing these tools are already embedded in teen social life.
I spent a little time with the blog last night and pulled together two new site features using Claude Cowork. The last time I experimented significantly with Claude like this was to use Claude chat to build the link log from scratch, walking it through my thinking in plain language, then copying and pasting its suggestions into the backend of the site and hitting publish. This time I used Cowork, the tool that runs in the browser, and it clicked through the screens itself, fully taking on the execution of tasks. I have some coding skills, but not the kind these changes required. If I were taking this on, I would need YouTube, Hugo for Dummies, and my own personal IT guy, and still probably couldn’t pull it together.
Last night I asked a few things of Claude:
I asked Claude Cowork to get into the backend of the micro.blog site and change my theme to link each line in the linklog back to my original post. The date field in the right column now links back to the original post for each link logged. No problem, easy request with easy execution. I made the request, confirmed the plan, and went about my business while Claude made the edits.
I also asked it to analyze my post content and suggest category tags for groups of content, then to label that content correctly in the backend. Claude reviewed about 500 posts and suggested I add a few new categories to my blog: AI in Practice, Books & Reading and Writing & Language. Then I set up some auto-filters to run at publish time to automatically categorize posts based on keywords moving forward.
Finally, I manually added the archive page, which lets you sort posts by category or year. This means I now have a functional archive here. Enjoy my anodyne thoughts, dear reader.
Adjacent to my day job, I’ve been toying with Claude Pro now for about a year; in my experience, it has improved significantly within the last six months. It’s not perfect: my requested edits were completed, but it also changed the CSS on the linklog so some of the text is too light to read, which I didn’t ask for and don’t want. But it has arguably extended my ability to execute on work that requires skills I don’t otherwise have (such as design and coding). What I do bring to the table is an expansive practical background in publishing and production, and all the language to describe it.