10114
Discover the future of AI with ChatGPT! 🤖💬 Get AI insights, tips, tricks, and more. Questions: @AI_inquires_bot
We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting.
If you’re an API developer trying to maximize performance, we recommend using the same settings that we deploy in our own products:
- Use our Responses API, not our legacy Chat
- Completions API
- Retain reasoning
- Use compaction
If you want to test your own mettle against frontier models, try the public games yourself at arcprize.org/tasks
We implemented the harness with the Responses API and turned on:
→ Retained reasoning
→ Context compaction
On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.
GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games?
We investigated. The harness was not letting it remember what it had learned.
We found that enabling two API settings tripled our scores with 6x fewer output tokens.
Install the open-source Codex Security CLI,:
npm install @OpenAI/codex-security
Or start with:
npx @OpenAI/codex-security@latest --help
NPM: npmjs.com/package/@opena…
At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone.
We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement.
We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible.
pacingthefrontier.com
Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verification, stewardship, and long-term maintenance matter.
Читать полностью…
Apply here:
https://openai.com/student-collective/
At a small business, the person closest to a problem is often the one who has to solve it—even when it falls outside their job description.
We studied how AI is becoming powerful generalist tool for small teams, helping them work across functions.
These 10 prompts turn it into a grocery-budget analyst, flight finder, personal trainer, bill negotiator, resume coach, study guide builder, travel planner, and more.
Читать полностью…
The statement calls for safety measures and stronger oversight before AI development moves beyond human control.
Читать полностью…
AI researchers are asking governments to slow down.
More than 1,100 employees from OpenAI, Anthropic, Google, Meta and other leading AI companies have signed “Pacing the Frontier,” a statement calling for international coordination to slow frontier AI development when necessary.
This job listing broke basic math
A Yahoo listing for a Senior Software Developer described Claude Code as a required skill, with more than 10 years of experience.
The AI proposed improving education and healthcare, modernising infrastructure, supporting clean energy, and making government more transparent. Some users praised the response as practical, while others stressed that ChatGPT generates answers from patterns in data, not personal beliefs.
Читать полностью…
A viral prompt asked ChatGPT how it would govern if it became President.
Читать полностью…
A benchmark score reflects the model as well as the harness and settings used to run it.
For long-running agents, retaining reasoning and compacting context lets the model build on what it has already learned.
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
ARC-AGI-3 tests how well models can learn unfamiliar 2D games without instructions.
The standard harness discarded GPT-5.6 Sol’s reasoning after each move and dropped earlier actions as the context filled up. The model had to keep starting over.
After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.
The results:
- 20% lower serving costs from production GPU kernel improvements.
- 15%+ better token-generation efficiency from improved speculative decoding.
These optimizations across our stack compound to unlock the most performant models at every point in the cost-intelligence curve.
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here...
You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD.
This is an early release, and we're listening to your feedback as we continue improving it.
🚨BREAKING: OpenAI says their unreleased model called 'Astra' cracked 10 math problems no human could solve for decades.
The problems span quantum complexity, post-quantum cryptography and group theory.
All ten proofs were formally verified and cost roughly $2,000 in tokens.
Coding agents are helping scientists spend more time advancing research, taking on everything from routine maintenance and targeted optimization to complete redesigns and new systems.
While agents can reliably execute on ambitious projects, researchers must still define the scientific questions, verify results, and take a stance on long-term ownership.
We're opening applications for the OpenAI Student Collective—a program for undergraduate Campus Leads bringing AI innovation to their campuses.
Campus Leads will work directly with OpenAI and receive hands-on training, funding, credits, swag, and access to a global community of peers.
Copy the prompt. Replace the brackets. Paste it into ChatGPT.
Читать полностью…
Most people type one vague sentence, get a vague answer, then decide ChatGPT is overrated.
The prompt is the difference.
The group warns that future AI systems could help automate AI research itself, accelerating progress faster than society can understand or control it. They argue that companies and countries may be unable to slow down independently because of competitive pressure.
Читать полностью…
The problem is that Claude Code only entered public preview in February 2025. Even Boris Cherny, the Anthropic engineer who built it, would not qualify for the role.
It is a perfect example of how quickly AI job requirements are becoming detached from reality.
The exchange shows how AI is becoming part of public conversations about leadership and government.
Читать полностью…
Its answer focused less on political personalities and more on using evidence to guide public decisions.
Читать полностью…