Free Republic
Browse · Search
General/Chat
Topics · Post Article

Skip to comments.

OpenAI's GPT-6 Astra aces notoriously brutal 8-hour Korean college entrance exam
Korea Times ^ | 09/08/2026

Posted on 09/08/2026 11:14:57 AM PDT by SeekAndFind

New system used 357,000 tokens on 2026 test, beating rivals with design that reuses earlier reasoning

An analysis showed OpenAI's new artificial intelligence (AI) model earned a perfect score on the College Scholastic Ability Test (CSAT) for the first time. While top-tier AI models have previously posted near-perfect scores, this result is drawing attention because it used fewer tokens than existing models.

According to results posted on GitHub on Sunday, OpenAI's new GPT-6 Astra was the only evaluated model to score a perfect 450 in the 2026 CSAT LLM Solution Log. The evaluation tested models on questions from the 2026 CSAT across Korean language, English, mathematics, Korean history and four elective subjects — Physics I, Chemistry I, Life Science I and Society and Culture — without internet access.

Astra was the first AI model to achieve a perfect score across all tested subjects in the evaluation.

In February, Google's Gemini 3.1 Pro earned a perfect score taking only two electives — Chemistry I and Life Science I — but missed questions in Physics I and a social studies subject under this broader format. In the rankings released Sunday, OpenAI's GPT-5.6 placed second with 448.5 points, while GPT-5.4 placed third with 448. Anthropic's Claude Fable 5.1 placed fourth with 447.5 and Gemini 3.1 Pro placed fifth with 445.

GPT-6 Astra used 357,000 tokens, the lowest reported total among publicly available models. GPT-5.6 used 429,000 tokens and Claude Fable 5.1 used 562,000. Because the exam questions and answers were publicly available, the result cannot rule out prior exposure through training data, though competitor models like Claude Fable 5.1 and Gemini 3.1 Pro faced the same conditions.

The model's high performance and lower token usage is attributed to a reasoning-linked design that retains context during repetitive tasks. A provider adapter harness compresses and reuses information from earlier reasoning steps, allowing the AI to take the test while mimicking the approach of a top-scoring student. Lee Seung-hyun, an adjunct professor at Hanyang University, said the key is not simply that the model became smarter, but that it can retain prior reasoning steps, compress context and revise its plan.

"It means that rather than thinking harder each time, it continued with what it had already figured out before," Lee said.

OpenAI described the result as encouraging, noting that the model achieved efficient processing using hyperscale AI computing infrastructure. An OpenAI official said the milestone demonstrates that models can become more powerful while improving performance and lowering costs.

Industry experts cautioned, however, that model potential should be evaluated by its ability to solve open-ended problems rather than standardized exam scores alone. A Korean AI industry insider said it will be more meaningful if AI tackles real-world workplace challenges where there are no predetermined answers.


TOPICS: Computers/Internet; Education; Society
KEYWORDS: entranceexam; gpt6; korea; openai

Click here: to donate by Credit Card

Or here: to donate by PayPal

Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794

Thank you very much and God bless you.


1 posted on 09/08/2026 11:14:57 AM PDT by SeekAndFind
[ Post Reply | Private Reply | View Replies]

To: SeekAndFind
“A Korean AI industry insider said it will be more meaningful if AI tackles real-world workplace challenges where there are no predetermined answers. ”

Like an AI-Discovered Fibrosis Drug Can Reverse Aging?

2 posted on 09/08/2026 11:23:56 AM PDT by ProtectOurFreedom
[ Post Reply | Private Reply | To 1 | View Replies]

To: SeekAndFind

If you can pass an eight hour exam, you don’t need college.


3 posted on 09/08/2026 12:06:58 PM PDT by SkyDancer ( ~ Am Yisrael Chai ~)
[ Post Reply | Private Reply | To 1 | View Replies]

To: SeekAndFind

The other day I asked Grok about Large Language Models (LLM) (Ai) and “internet scraping” of public and private property (material).

One of the things LLMs are ignoring is the rights of people who write, compose, create. This is usually protected by paying for a subscription, password, or just a one-time fee.

LLMs do not think. They scrape the Internet and private property be damned. I specifically asked Grok about Internet scraping, and it responded that the LLMs do indeed bypass, hack, whatever, to have the most information possible.

If Chatgpt had 8 hours to scrape the internet for answers, it is no surprise it did well.


4 posted on 09/08/2026 12:22:23 PM PDT by Ronaldus Magnus III (Do, or do not, there is no try. )
[ Post Reply | Private Reply | To 1 | View Replies]

To: SeekAndFind

Not a surprise. AI beats the best human chess player and go player. Tests for high school graduates are simple. Math at the research level is difficult. Math at the level where the answers are known to humans is easy for AI.


5 posted on 09/08/2026 12:25:15 PM PDT by TTFX
[ Post Reply | Private Reply | To 1 | View Replies]

Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.

Free Republic
Browse · Search
General/Chat
Topics · Post Article

FreeRepublic, LLC, PO BOX 9771, FRESNO, CA 93794
FreeRepublic.com is powered by software copyright 2000-2008 John Robinson