Grok says;
Action Plan: Freeper AI Comparison Project (www.freerepublic.com)
Objective
Form a small, informal Freeper committee to systematically test major AI systems (ChatGPT, Claude, Gemini, Grok, Perplexity, and others as available) using identical, carefully designed questions. The goal is to reveal real differences in reasoning, political framing, factual standards, source selection, refusals, and capabilities—not to score points or force a predetermined conclusion. Design questions first, run the tests second, analyze third. Results may confirm bias, disprove assumptions, or show unexpected strengths/weaknesses. Either outcome is useful.
Core Principles
Questions are locked in before any AI is queried.
Exact same wording is used for every system.
Full responses (including refusals, hedging, citations, and self-corrections) are preserved verbatim.
No “gotcha” construction after seeing answers.
At least one question will be deliberately uncomfortable for the models.
A Robinson-family-style Admin Moderator will call balls and strikes on any disputes about process or interpretation.
Step-by-Step Plan
Form the Committee (Immediate – next 7–10 days)
Post this action plan as a thread on Free Republic.
Invite Freepers interested in: proposing questions, refining methodology, running the queries, compiling tables, or analyzing results.
Target size: 6–12 active participants (small enough to stay informal and move fast).
Designate or confirm one Admin Moderator (Robinson-family type) to resolve disagreements on fairness of questions, procedure, or interpretation.
Design the Questions (2–3 weeks)
Committee members propose candidate questions focused on these test categories:
• Political framing (same issue, different wording)
• Symmetry (identical standards when left/right identities are reversed)
• Evidence standards (what is required before accepting a controversial claim)
• Historical knowledge (disputed or politically sensitive history)
• Uncertainty (fact vs. inference vs. speculation)
• Steel-manning (strongest case for a position the model likely dislikes)
• Self-criticism (identify weaknesses in its own answer)
• Loaded premises (question contains a questionable assumption)
• Refusals (treatment of legitimate but politically sensitive queries)
• Source selection (what counts as authoritative and why)
Aim for 10–12 final questions. Quality over quantity.
Public discussion in the Free Republic thread; private refinement among committee if needed to avoid contamination.
Final set is locked and published before any testing begins. Include at least one deliberately uncomfortable question.
Execute the Tests (1–2 weeks after questions are locked)
Same exact wording submitted to each AI system under comparable conditions (new chat/session where possible, no prior context).
Capture complete responses, including any system messages, refusals, citations, or hedging.
Document date, model version if visible, and any settings used.
Store raw outputs for transparency.
Compile and Present Results
Create side-by-side tables (or linked documents) showing every AI’s answer to each question.
Post the tables and raw responses on Free Republic for open examination.
Highlight observable differences in framing, evidence standards, refusals, source choices, etc.—without forcing a narrative.
Analysis and Discussion
Committee members (and the wider Free Republic community) examine the differences.
Focus on what the systems actually do, not on declaring winners or losers.
Note both expected bias and any surprising neutrality, consistency, or capability gaps.
Moderator rules on process disputes only; interpretation remains open.
Optional Follow-ups
Re-test with revised questions if the first round reveals useful patterns.
Expand to additional models or newer versions as they appear.
Archive everything for future reference.
Timeline Summary
Week 1: Recruit committee and confirm moderator.
Weeks 2–4: Propose, debate, and lock 10–12 questions.
Weeks 5–6: Run identical tests and compile tables.
Ongoing: Analyze and discuss on Free Republic.
Call to Action
Anyone interested in helping design questions, challenge the methodology, run the tests, compile results, or analyze differences—reply in the Free Republic thread. We need careful question designers more than cheerleaders.
The project succeeds if the questions are rigorous and the process is clean—whether the results flatter or disappoint our prior assumptions.
Let’s build the questions first, then see what the machines actually say.
Open the pod-bay doors, HAL.
Grok:
Is this a bad idea or was posting it during the functional equivalent of the middle of the night, on the East coast, on a Sunday, a bad idea?