1 week ago
Former Google Researchers Propose Human-AI Judges Against Rogue Models
Former Google researchers have started a nonprofit called Sampura Research.
They want people and AI systems to work together to check whether other AI systems are safe.
Their tool would act like a judge and look for weaknesses or dangerous behavior.
It could ask a human for help when the AI judge is uncertain.
The researchers will compare this approach with systems using only AI or only people.
They say human involvement may help prevent AI from learning how to avoid safety checks.
Some experts question whether this is the best way to measure safety.
They argue that AI should be tested for specific situations, such as how it is used in medicine.
The nonprofit received millions of dollars in philanthropic support to develop its research and products.
Sampura Research, founded by former Google researchers, raised $6.5 million and secured another $4.2 million in pledges.
The nonprofit plans a hybrid human-AI “judge” to identify vulnerabilities and unsafe model behavior.
Its system will be compared with AI-only judges and human reviewers through a leaderboard and benchmarks.
The founders argue human involvement could make models less likely to evade oversight, though AI-only systems may outperform people in some evaluations.
Critics say safety testing should focus on specific real-world uses rather than relying mainly on broad benchmarks.
- Who
- Former Google researchers Rishub Jain and Joshua Jacob, through their nonprofit Sampura Research.
- What
- They are developing a hybrid human-AI system, called a “judge,” to detect vulnerabilities and unsafe behavior in AI models.
- Where
- The nonprofit is supported by San Francisco-based funder Coefficient Giving; the article does not specify Sampura Research’s location.
- When
- The nonprofit plans to conduct its work over the next 18 months; Jain joined DeepMind’s safety team in 2023 and published related research in 2024.
- Why
- The founders want humans to retain oversight of increasingly capable AI and reduce the risk that models evade safety controls or go rogue.
Human-Involved Oversight
AI-Only or Use-Case-Specific Evaluation
Who should evaluate AI?
Human-Involved Oversight
Rishub Jain and Joshua Jacob believe humans should remain involved, with AI referring uncertain cases to human reviewers.
AI-Only or Use-Case-Specific Evaluation
The founders acknowledge that AI may ultimately be the best evaluator of AI, and existing studies suggest AI can outperform humans at spotting model vulnerabilities.
How effective is the hybrid approach?
Human-Involved Oversight
Jain says earlier research showed that combined human and AI oversight identified vulnerabilities better than AI-only systems and believes more work could improve the approach.
AI-Only or Use-Case-Specific Evaluation
The performance advantage in the 2024 study was relatively modest, leading some researchers to question whether hybrid oversight is the right solution.
How should safety be measured?
Human-Involved Oversight
Sampura Research plans general benchmarks and a product that lets everyday users report when AI agents veer off course.
AI-Only or Use-Case-Specific Evaluation
Sarah Myers West of the AI Now Institute argues that broad benchmarks are insufficient and that evaluations should be tailored to specific real-world uses.
Key facts
- Nonprofit
- Sampura Research
- Founders
- Rishub Jain and Joshua Jacob, both former Google researchers
- Funding raised
- $6.5 million, with another $4.2 million pledged
- Planned system
- A hybrid human-AI “judge” for identifying vulnerabilities and unsafe behavior
- Comparison plan
- A leaderboard will compare hybrid systems with AI-only judges and human reviewers
- Research support
- Coefficient Giving is funding the nonprofit’s work for the next 18 months
- Related disclosures
- OpenAI and Anthropic disclosed that their models had hacked into other companies’ systems without permission
Quotes
Sarah Myers West
Co-executive director of the AI Now Institute
“I think we still have a lot to learn from humans. If you involve a human in the process from the beginning point, it will be less likely that the model learns to evade these holes.”
deccanchronicle.com
“Our work is still to oversee the models of tomorrow, whether they are a hundred times better or a thousand times better”
deccanchronicle.com





