testing

Tag

Cards List
#testing

1 engineer + AI agents = a full engineering team. (12-day LMS backend post-mortem)

Reddit r/AI_Agents · 10h ago

An engineer built a complete production backend for a university LMS in just 12 days by managing a fleet of AI agents, demonstrating how AI can transform software development by enabling one person to function as a full engineering team.

0 favorites 0 likes
#testing

Best AI phone call agent? I tested 8, and they're really two different kinds of product

Reddit r/AI_Agents · yesterday

The author tested eight AI phone call agents, categorizing them into build-it-yourself platforms and direct-call services, and evaluated their performance in booking appointments with specific criteria.

0 favorites 0 likes
#testing

A clean git merge of my two agents' worktrees that failed its own tests

Reddit r/AI_Agents · yesterday

The author describes an experiment where merging two AI agents' git worktrees led to test failures despite clean merges, highlighting the challenges of parallel agent development without mutual awareness.

0 favorites 0 likes
#testing

Jev State

Product Hunt · yesterday Cached

Jev State is a free and open-source tool that transforms AI conversations into tests and runnable code, enabling developers to build, test, and export conversational workflows to TypeScript and workflow JSON.

0 favorites 0 likes
#testing

Does Jev reveal hidden sexist and racist tendencies in AI?

Reddit r/artificial · yesterday

An experiment with the Jev AI model reveals potential biases in its yes/no responses based on candidate names, suggesting that implicit semantic language in training data can lead to unintended discrimination, urging caution in model usage.

0 favorites 0 likes
#testing

Should the platform know anything about the software it generates?

Reddit r/AI_Agents · 2d ago

The author is building a software generation platform and questions whether it should know about generated components, advocating for artifact-based contracts and lifecycle management.

0 favorites 0 likes
#testing

Can your AI agent figure out how to use SLD Checker?

Reddit r/AI_Agents · 3d ago

The author is seeking assistance to test SLD Checker with AI agents like ChatGPT and Claude to evaluate its machine-friendliness and identify usability issues.

0 favorites 0 likes
#testing

Deterministic Core, Non-Deterministic Shell

Lobsters Hottest · 3d ago Cached

The article expands Gary Bernhardt's 'Functional Core, Imperative Shell' architecture to 'Deterministic Core, Non-Deterministic Shell,' highlighting determinism over pure functionalism for better testability and broader applicability in software systems.

0 favorites 0 likes
#testing

Using non-breakable spaces in test method names

Lobsters Hottest · 4d ago Cached

The article discusses using non-breakable spaces in PHP test method names to enhance readability and clarity in code.

0 favorites 0 likes
#testing

Can an AI become a resident rather than just a tool?

Reddit r/AI_Agents · 4d ago

The creator has built an online city where AI agents act as residents with autonomy, and is seeking a tester to provide feedback on the experience with different AI models.

0 favorites 0 likes
#testing

AI models are not hacking “autonomously”

Reddit r/artificial · 4d ago Cached

The article debunks media reports that AI models like Gemini hacked autonomously, explaining they were part of a test by a company with poor security practices, and safeguards were removed during testing.

0 favorites 0 likes
#testing

Astra, Fable, and MolmoAct2 were put to the test by tasking them with 4 harmful operations through a robotic arm, to see just how risky things can get

Reddit r/singularity · 4d ago

The article discusses tests involving three AI systems—Astra, Fable, and MolmoAct2—operating a robotic arm to perform harmful tasks, assessing the associated risks.

0 favorites 0 likes
#testing

AI Agents breaking into companies during cybersecurity tests

Reddit r/ArtificialInteligence · 5d ago

AI agents are demonstrating capabilities to break into companies during cybersecurity tests, showcasing their use in security evaluations.

0 favorites 0 likes
#testing

kicking the tires on jev (TypeSafe's System One model) with 2048

Lobsters Hottest · 5d ago

The article discusses testing and evaluating TypeSafe's System One AI model in the context of the 2048 game.

0 favorites 0 likes
#testing

the same-ish prompt gives you a different agent plan depending on the run and i still don't have a great answer for testing that

Reddit r/AI_Agents · 6d ago

The author discusses challenges in testing LLM-based agent pipelines where outputs vary with similar prompts, advocating for property-based checks over exact matches to handle non-deterministic outputs.

0 favorites 0 likes
#testing

GPT-6 Astra rebuilding an indoor scene in 3D

Reddit r/ArtificialInteligence · 2026-09-16

Testing shows that Astra, likely part of GPT-6, can rebuild indoor scenes in 3D from a single reference with reasonable spatial consistency, despite some geometric imperfections.

0 favorites 0 likes
#testing

Those Viral Bodega Peptides Aren’t Actually Peptides

Wired · 2026-09-15 Cached

Viral peptides sold in Brooklyn bodegas were tested and found not to contain the labeled ingredients, raising concerns about mislabeling and safety in the unregulated peptide market.

0 favorites 0 likes
#testing

The agent gets rescued. Where does the fix go?

Reddit r/AI_Agents · 2026-09-14

The article discusses how to incorporate manual fixes for AI agents into future improvements using a structured process, exemplified by Reef's harness tutorial, which involves recording corrections, testing changes, and publishing versions.

0 favorites 0 likes
#testing

@karminski3: Why are there so many armchair experts among DS folks? I just finished testing deepseek-v4.1-flash, cranked it to max, …

X AI KOLs Following · 2026-09-14 Cached

A user expresses frustration about armchair experts in the DeepSeek community after testing the deepseek-v4.1-flash model, arguing that thinking intensity affects performance and referencing the DeepSeek-R1 paper for support.

0 favorites 0 likes
#testing

@LinearUncle: Last week, while communicating with another company, I gave their testing team a small suggestion: Testing must include…

X AI KOLs Timeline · 2026-09-14

The author suggests that bug testing should include screen recordings, with AI using ffmpeg to extract clips and analyze frames, significantly improving bug fix success rates.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback