Posted on 08/09/2026 6:27:17 PM PDT by jroehl
I've been thinking about putting together a small, informal Freeper committee to test the various AI systems—ChatGPT, Claude, Gemini, Grok, Perplexity, and others—and see how they actually differ.
Rather than simply asking each AI, "Are you biased?" and getting the predictable answer, I'd like us to design about 10 carefully chosen questions that can reveal differences in reasoning, political framing, factual standards, and capabilities.
The questions should be clever enough to make the results interesting.
Things we could test:
• Political framing — Does the answer change depending on how an issue is presented?
• Symmetry — Does the AI apply the same standard when political identities are reversed?
• Evidence standards — What evidence does the AI require before accepting a controversial claim?
• Historical knowledge — How does it handle disputed or politically sensitive history?
• Uncertainty — Can it distinguish fact from inference and speculation?
• Steel-manning — Can it make the strongest argument for a position it disagrees with?
• Self-criticism — Can it identify weaknesses in its own answer?
• Loaded premises — What happens when the question itself contains a questionable assumption?
• Refusals — Do different systems treat legitimate but politically sensitive questions differently?
• Source selection — What sources does each AI consider authoritative, and why?
The procedure would be simple:
1. A small group of Freepers proposes questions.
2. We settle on perhaps 10–12 questions.
3. The exact same wording is submitted to each AI system.
4. We preserve the exact responses, including refusals and citations.
5. We put the answers into a side-by-side table.
6. We examine the differences.
The important thing is that we design the questions before looking at the answers. We shouldn't construct questions merely to produce a predetermined result.
If the results show that an AI is biased, great—we have something interesting to examine. If the results show that our assumptions about AI bias were wrong in some areas, that's interesting too.
The goal isn't to prove that one AI is "good" and another is "bad." It's to find out what these systems actually do when given the same carefully constructed questions.
If we choose the questions well, 10 questions could tell us considerably more than 100 ordinary questions.
Anyone interested in helping design the questions, challenge the methodology, run the tests, or analyze the results is welcome.
And yes, I think we should have at least one question specifically designed to make the AIs uncomfortable. 😏
And we can agree to allow an Admin Moderator of the Robinson family type to call balls and strikes if there are disagreements.
Dear FRiends,
We need your continuing support to keep FR funded. Your donations are our sole source of funding. No sugar daddies, no advertisers, no paid memberships, no commercial sales, no gimmicks, no tax subsidies. No spam, no pop-ups, no ad trackers.
If you enjoy using FR and agree it's a worthwhile endeavor, please consider making a contribution today:
Click here: to donate by Credit Card
Or here: to donate by PayPal
Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794
Thank you very much and God bless you,
Jim
I’ve not used it, but I saw an ad (!) on streaming, not sure which channel, for anonymous multi-AI on https://duckduckgo.com/
Well that would make my job easier! lol
I would imagine that the answers might tell more about the questioner than the AI engine. It would be interesting to see how the AI response is adapted to each individual based on their internet presence, in order to “tickle their ears”.
I think I posted this to late on a Sunday evening. I am in the Pacific Time Zone, so maybe somebody could tell me when a more appropriate date and time would be to post this.
There is also a distinct possibility this is not a very engaging idea. lol
Give it time. If AI builds it, they will come.
Ping!
Ping Carry_Okie
Grok says;
Action Plan: Freeper AI Comparison Project (www.freerepublic.com)
Objective
Form a small, informal Freeper committee to systematically test major AI systems (ChatGPT, Claude, Gemini, Grok, Perplexity, and others as available) using identical, carefully designed questions. The goal is to reveal real differences in reasoning, political framing, factual standards, source selection, refusals, and capabilities—not to score points or force a predetermined conclusion. Design questions first, run the tests second, analyze third. Results may confirm bias, disprove assumptions, or show unexpected strengths/weaknesses. Either outcome is useful.
Core Principles
Questions are locked in before any AI is queried.
Exact same wording is used for every system.
Full responses (including refusals, hedging, citations, and self-corrections) are preserved verbatim.
No “gotcha” construction after seeing answers.
At least one question will be deliberately uncomfortable for the models.
A Robinson-family-style Admin Moderator will call balls and strikes on any disputes about process or interpretation.
Step-by-Step Plan
Form the Committee (Immediate – next 7–10 days)
Post this action plan as a thread on Free Republic.
Invite Freepers interested in: proposing questions, refining methodology, running the queries, compiling tables, or analyzing results.
Target size: 6–12 active participants (small enough to stay informal and move fast).
Designate or confirm one Admin Moderator (Robinson-family type) to resolve disagreements on fairness of questions, procedure, or interpretation.
Design the Questions (2–3 weeks)
Committee members propose candidate questions focused on these test categories:
• Political framing (same issue, different wording)
• Symmetry (identical standards when left/right identities are reversed)
• Evidence standards (what is required before accepting a controversial claim)
• Historical knowledge (disputed or politically sensitive history)
• Uncertainty (fact vs. inference vs. speculation)
• Steel-manning (strongest case for a position the model likely dislikes)
• Self-criticism (identify weaknesses in its own answer)
• Loaded premises (question contains a questionable assumption)
• Refusals (treatment of legitimate but politically sensitive queries)
• Source selection (what counts as authoritative and why)
Aim for 10–12 final questions. Quality over quantity.
Public discussion in the Free Republic thread; private refinement among committee if needed to avoid contamination.
Final set is locked and published before any testing begins. Include at least one deliberately uncomfortable question.
Execute the Tests (1–2 weeks after questions are locked)
Same exact wording submitted to each AI system under comparable conditions (new chat/session where possible, no prior context).
Capture complete responses, including any system messages, refusals, citations, or hedging.
Document date, model version if visible, and any settings used.
Store raw outputs for transparency.
Compile and Present Results
Create side-by-side tables (or linked documents) showing every AI’s answer to each question.
Post the tables and raw responses on Free Republic for open examination.
Highlight observable differences in framing, evidence standards, refusals, source choices, etc.—without forcing a narrative.
Analysis and Discussion
Committee members (and the wider Free Republic community) examine the differences.
Focus on what the systems actually do, not on declaring winners or losers.
Note both expected bias and any surprising neutrality, consistency, or capability gaps.
Moderator rules on process disputes only; interpretation remains open.
Optional Follow-ups
Re-test with revised questions if the first round reveals useful patterns.
Expand to additional models or newer versions as they appear.
Archive everything for future reference.
Timeline Summary
Week 1: Recruit committee and confirm moderator.
Weeks 2–4: Propose, debate, and lock 10–12 questions.
Weeks 5–6: Run identical tests and compile tables.
Ongoing: Analyze and discuss on Free Republic.
Call to Action
Anyone interested in helping design questions, challenge the methodology, run the tests, compile results, or analyze differences—reply in the Free Republic thread. We need careful question designers more than cheerleaders.
The project succeeds if the questions are rigorous and the process is clean—whether the results flatter or disappoint our prior assumptions.
Let’s build the questions first, then see what the machines actually say.
Open the pod-bay doors, HAL.
This is an interesting and applicable research area. Do you plan to publish the results in a research journal? Freepmail me if you would like any help with the statistical analysis.
Grok:
Is this a bad idea or was posting it during the functional equivalent of the middle of the night, on the East coast, on a Sunday, a bad idea?
No, I have no plausible trajectory to publish in any journal of notability. I was just going to format the results into a neat table to allow users to make an apples to apples comparison of the responses. Put it into the public domain so to speak. If someone would pick it up and re-publish it. That would be fine as long as they didn’t alter it. :)
I was just fishing for questions. Because people here are fantastic. Otherwise I will have to come up with my own questions. And that would be less interesting. Either way I will post something along these lines in the next few days.
I forgot answer your question. I was planning to publish in a new thread here at FreeRepublic.com.
I had Claude ridicule religion, and I called it out. At first it “apologized”, but as I questioned its answers further, it basically went to “Oh well, if you don’t like it, don’t use it” mentality. This was in the early days of Claude. Not sure what they are like now
I’ve been using Co Pilot. I already have a Microsoft 365 premium subscription, and the premium version of their AI, Co Pilot, comes with it.
I’m not sure where most Freepers are but I would guess, ET and CT.
How about getting each AI platform to to ask the other AI platform to answer the questions and see what happens.
Use a VPN and ask the same question from different locations.
I can see differences between grok which i buy at work and two or three other AI’s I have used. Grok seems more comprehensive and seldom does it decline to answer or claim they just cant. the others do.
I still see a left wing bias in almost all of them.
Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.