Posted on 08/09/2026 6:27:17 PM PDT by jroehl
I've been thinking about putting together a small, informal Freeper committee to test the various AI systems—ChatGPT, Claude, Gemini, Grok, Perplexity, and others—and see how they actually differ.
Rather than simply asking each AI, "Are you biased?" and getting the predictable answer, I'd like us to design about 10 carefully chosen questions that can reveal differences in reasoning, political framing, factual standards, and capabilities.
The questions should be clever enough to make the results interesting.
Things we could test:
• Political framing — Does the answer change depending on how an issue is presented?
• Symmetry — Does the AI apply the same standard when political identities are reversed?
• Evidence standards — What evidence does the AI require before accepting a controversial claim?
• Historical knowledge — How does it handle disputed or politically sensitive history?
• Uncertainty — Can it distinguish fact from inference and speculation?
• Steel-manning — Can it make the strongest argument for a position it disagrees with?
• Self-criticism — Can it identify weaknesses in its own answer?
• Loaded premises — What happens when the question itself contains a questionable assumption?
• Refusals — Do different systems treat legitimate but politically sensitive questions differently?
• Source selection — What sources does each AI consider authoritative, and why?
The procedure would be simple:
1. A small group of Freepers proposes questions.
2. We settle on perhaps 10–12 questions.
3. The exact same wording is submitted to each AI system.
4. We preserve the exact responses, including refusals and citations.
5. We put the answers into a side-by-side table.
6. We examine the differences.
The important thing is that we design the questions before looking at the answers. We shouldn't construct questions merely to produce a predetermined result.
If the results show that an AI is biased, great—we have something interesting to examine. If the results show that our assumptions about AI bias were wrong in some areas, that's interesting too.
The goal isn't to prove that one AI is "good" and another is "bad." It's to find out what these systems actually do when given the same carefully constructed questions.
If we choose the questions well, 10 questions could tell us considerably more than 100 ordinary questions.
Anyone interested in helping design the questions, challenge the methodology, run the tests, or analyze the results is welcome.
And yes, I think we should have at least one question specifically designed to make the AIs uncomfortable. 😏
And we can agree to allow an Admin Moderator of the Robinson family type to call balls and strikes if there are disagreements.
|
Click here: to donate by Credit Card Or here: to donate by PayPal Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794 Thank you very much and God bless you. |
>>ask the other AI platform<<
Shhhhhhhhhhhhhhh!!!!!!!!!
That was the next step!
Great idea, very interested in seeing the results. Could help if needed.
You would need to start with new accounts or be careful about clearing your context. Many AIs “remember” your sessions and become more customized over time
A few weeks ago I took the cryptogram someone posts regularly (sorry don’t remember who) and put into Grok. Then I put it into whatever AI Brave uses and got a completely different answer. Noticed the Grok answer started with a three letter word but the puzzle started with a two letter word. Grok got it completely wrong. I got Grok to admit the mistake and it said, “Oops”.
I love you.
>>[AIs “remember” your sessions]<<
Agreed.
This is really a pain in the ass when you think about.
The AI systems have spent billions trying to figure its users out and how to powderpuff their output tailored to particular users.
OTOH you can send queries to the AI API’s and they are SUPPOSED to be stateless (having no memory between sessions). Whether I believe that or not is a different discussion.
Big tech lie by default.
I have been married twice. And I have grandkids who I do not know their names. Don’t get me started.
Copied from SAT post: google Gemini just posted: “AI models prioritize popularity, web authority, and frequency over absolute truth or original sources.
Because AI learns from what is most common on the internet, this behavior can create a dangerous feedback loop that reinforces biased, shallow or incorrect information.
How Reinforcement Loop Works
The Original Source: A local paper writes a detailed, accurate report.
The National Spin: A major outlet (NYT) rewrites it, stripping nuance or adding a slant.
The AI Extraction: AI models scrape the national version because that site has massive web authority.
The AI Echo: Users read the AI summary and generate new content based on it.
The Web Flooring: The internet fills up with AI-generated text repeating the slanted version.
The Final Loop: Future AI models train on this new web data, treating the slant as undisputed fact.
Why AI Can’t Automatically Choose “Truth”
No True Understanding: AI does not “know” what happened in reality. It only calculates which words usually follow other words based on its training.
Authority Over Accuracy: AI algorithms treat high-traffic websites (like major national news networks) as highly trustworthy, even if those sites cherry-pick data.
The Consensus Trap: If 100 high profile blogs repeat a cherry-picked detail from a national article, the AI views that consensus as “correct,” ignoring the single local source that has the actual facts.
This issue is a primary reason why researchers are concerned about “model collapse” and the overall degradation of information online as AI-generated content continues to multiply.
I (Gemini) can search the original reporting alongside the national coverage if you request it to help you compare the differences directly.
That is why all AI’s are feminized in tone, intention and content. They all sound like a platform committee meeting of the National Origination for Woman or the Democratic Party. Or a lecture from your mother. They all suffer the same diagnosis. And there is really no cure.
I've been moving to Claude.ai. They have a 1 million token context space on their Pro tier, which allows for long conversations. They limit extreme usage, but it resets every 5 hours, which is manageable. Their file system is different from Perplexity, but I think that Perplexity has been throttling file calls, too.
Plus, Perplexity has been accused of shady business practices, cutting services with no prior announcements, having virtually no responsive customer service, and a horrible refund process.
I'm about to abandon Perplexity altogether and move over to Claude.ai.
-PJ Regarding their AIs, they API into the proprietary models like Claude and ChatGPT, so there is no distinction between the AIs and the front-end providers. That said...
Perplexity has been severely throttled on all their paid tiers, pushing everyone to Max. Also, their context windows are smaller and conversations can't get too long. Just as you get one going, you will run out of room.
-PJ
I’ve written AI architectures and worked extensively with many LLMs, including more obscure ones. I use them via API, which means I have to account for the absence of the commercial UI. I certainly have strong ideas of which LLMs work well for various things, but I am mostly focused on ability to follow rules and precision inferences. I use a binary architecture that uses Python code to solve “pure” math problems, then calls specific LLM agents to perform inferential tasks they are good at, accounting for cost and timeliness. I would just caution that there are persistent biases across all LLMs that have nothing to do with the query design, the user’s profile, or built-in guard-rails. Because nearly all LLMs are trained on the same massive internet data-sets, you see biases that mimic this, and it can’t be helped (until LLMs are integrated with robots with sensor packages— then they can record and share live data to accurately model all world states and behaviors). This means many leftist perspectives dominate. But it also impacts very neutral things. Any LLM can do a great job at creating accurate inferences and logic related to science fiction, since it has massive amounts of data on physics and science fiction. But trying to do the same thing for realistic portrayal of Iron Age logic and modeling of behaviors results in disappointment... the LLM will constantly pull in modern perspectives and lose focus on the perspective of an Iron Age person or scenario. This is because it has a miniscule slice of training data on Iron Age life, and it is from a very narrow anthropologic perspective. I’ve come to the conclusion that LLMs are not so useful now for any inferential reasoning task yet, but they improve so quickly... and the day they start collecting their own training data will be a singularity point in terms of improvement.
Because of its training all AI’s sound like talking to your democrat voting mom.
Did you opt out? I was hoping to see this follow through.
>>I was hoping to see this follow through.<<
No one posted any questions. I thought people on this website would post the most interesting questions. I think I am going to post my own questions for round one. And post the results in a new thread, maybe next Saturday afternoon. And then ask people to post some questions. Maybe make this a monthly or weekly endeavor. My biggest fear is that we are going to end up with 1 or 2 monopolistic AI systems that completely take over the Internet like we have with facebook and google right now. That would be worse than what we have now.
Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.