prompt-testing

Tag

Cards List
#prompt-testing

@AISuperDomain: OpenAI is specifically screwing over domestic users! It turns out that many Codex users who set the reasoning level to Ultra are actually assigned a downgraded gpt-5.5-mini model. But most users haven't noticed at all — they just mindlessly run loops. However, the detection method is also simple — you can use the prompt below to test whether it is...

X AI KOLs Timeline · 2026-08-10 Cached

The post claims that OpenAI Codex assigns a downgraded gpt-5.5-mini model to domestic users at the Ultra reasoning level, provides a detection prompt, and calls on users to reply 1 or 2 to measure the downgrade rate for domestic users.

0 favorites 0 likes
#prompt-testing

smevals - a small eval suite for evaluating models, prompts, and harnesses

Simon Willison's Blog · 2026-07-31 Cached

Simon Willison introduces smevals, a small eval suite from Prime Radiant for evaluating models, prompts, and harnesses, with commands to run evals, grade results, and serve static HTML reports.

0 favorites 0 likes
#prompt-testing

@ArizePhoenix: You can use PXI to run an experiment directly from Phoenix! Here's one that tests the system prompt vs. schema-aware pr…

X AI KOLs Following · 2026-07-27 Cached

Arize Phoenix demonstrates using PXI to run an experiment comparing system prompt vs schema-aware prompt with a programmatic code evaluator, avoiding the need for an LLM judge.

0 favorites 0 likes
#prompt-testing

We built an automated QA/eval engine for agent prompts. Help us test it out!

Reddit r/AI_Agents · 2026-07-10

Built an automated QA/eval engine for agent prompts called Baseline that treats prompts like software for regression testing, allowing non-coders to define rubrics and automatically optimize prompts. Currently in limited beta with a 30-day free trial.

0 favorites 0 likes
#prompt-testing

Natural-Language Testing for AI Agents (using simulated isolates)

Reddit r/AI_Agents · 2026-06-28

This article introduces a new natural-language testing system for AI agents that uses simulated isolates to automatically generate multi-turn simulations and evaluate agent behavior, helping developers catch regressions from prompt changes.

0 favorites 0 likes
#prompt-testing

Someone made my AI dream tool

Reddit r/artificial · 2026-06-02

A post highlights AIfiesta.ai, a tool that displays responses from multiple AI models (ChatGPT, Gemini, Claude) to the same prompt simultaneously, each in its own column.

0 favorites 0 likes
#prompt-testing

@punk2898: It's still Grok. Generated once, missed the second time. Prompt: "A cheerful young Chinese woman, long straight brown hair, natural makeup, wearing a soft beige spaghetti strap camisole, doing a plank on a blue mat in a bright gym studio with mirrors and equipment in the background, soft natural lighting, editorial photography style, photorealistic, ..."

X AI KOLs Timeline · 2026-05-17 Cached

User tested Grok's image generation function and found that the first time it successfully generated a complete image, but the second time it missed part of the prompt content, resulting in an incomplete generation.

0 favorites 0 likes
← Back to home

Submit Feedback