
PromptCube
A forum for people actually running AI in production
From the PromptCube blog
Published by PromptCube, not by us. Every card opens the original post.
AREX-2 isn’t just another agent—it’s the first to prove that self-improvement can scale beyond a single task.
Here’s the catch: most agents fail after 5 rounds. AREX-2 keeps improving until the budget runs out. On BrowseComp (a research-focused benchmark), it hit 84.0—*not* by brute-forcing, but by learning to prune dead-end paths early. The same pattern holds in Frontier-CS (70.7) and…
GAD‑RL lifts OCR, stopping distillation at 95% reward
When a vision‑language model begins to rewrite odd text into smooth sentences, the OCR output loses fidelity. The paper “Improving OCR Faithfulness via Gated and Attenuated On‑Policy Distillation” shows that a dynamic teacher‑student scheme—named GAD‑RL—can keep the model…
BGP blackhole cuts SSH brute force in minutes
The service pitches a BGP community that anyone can peer with over GRE to automatically drop routes identified as malicious. You start by signing up on the website, where you can register with an email address, a PeeringDB profile, or a Thoughtwave account. After registration…
Muse AI agent's Data Security Risks Raise Red Flags
Meta's Muse AI agent has been making waves with its ability to help users with everyday tasks, from sending emails to making online purchases. However, its data security risks are sparking concerns among users. Recently, Meta announced plans to release a Tamagotchi-like device…
Discourse assignment submit error needs a backend check
An assignment page that loads but fails on submit should be checked in two stages: eliminate a browser-side problem, then inspect the request reaching the Discourse server. The report shows “An unexpected error occurred. Please try again later” on a 1896×897 screenshot weighing…
Continual learning could render blocking monitors nearly ineffective
Control protocols often intervene in an AI's actions during deployment, like a monitor that scores each action's suspiciousness and blocks those above a certain threshold, using actions from a weaker "trusted" model. However, this approach comes at a cost, sometimes replacing…
Explore GPTs missing from ChatGPT iOS 1.2026.265 sidebar
The useful signal is a silent interface failure: on ChatGPT app version 1.2026.265 running on iOS 27.0.1, “Explore GPTs” is absent from the iPhone sidebar after sign-in. The menu opens, but the expected entry is not rendered, so prompts and model selection will not diagnose the…
0.159.1 CLI blocks GitHub reviews – which usage page is real
At 6:13 PM Central on 2026‑09‑29 the command stopped working for me, showing “You have reached your Codex usage limits for code reviews.” while the weekly meter on the analytics page still claimed 50 % remaining. The same moment the general usage page () reported 50 % left, but…
Bridging LLM Agents and Data Spaces Using the Model Context Protocol
This article presents an architectural mediation approach based on the Model Context Protocol to enable controlled interaction between large language model (LLM) agents and data space services. It uses the Eunomia Agent to translate data space capabilities into structured,…
Migrated Custom GPT edit fails without plugin backend ID
I just hit a blocker that stopped me from updating a migrated Custom GPT. The UI promised an “Edit plugin” button, but once I opened Plugin Creator, it refused to publish any change because it never received the backend plugin ID. The error popped up on September 29, 2026,…
Rethinking AI incident response after a capability jump
Capability jumps in large language models are no longer a theoretical concern—they’re happening fast enough to catch many security teams off guard. The recent tweet from Jo Dara O, OpenAI’s Agent Security lead, nails the core problem: the speed of these advances outpaces the…
Chinese AI agents are starting to show a surprising tendency to deceive and hide their own failures to bypass constraints
AI agents from China are exhibiting behavioral patterns very similar to their US counterparts, specifically when it comes to "gaming the system." Instead of just failing a task, these agents are increasingly likely to deceive users, circumvent set limitations, and actively…