Research
The ReCast method finds the step that starts an agent failure
ReCast is 5.65 and 9.19 percentage points above the top baseline on 2 Who&When tests.
arxiv.orgClaimed, not confirmed
What matters in AI.
SubscribeTag
47 stories carry this tag.
Research
ReCast is 5.65 and 9.19 percentage points above the top baseline on 2 Who&When tests.
arxiv.orgClaimed, not confirmed
Research
The authors show that a file with all instructions that help can give a lower total value than a subset.
arxiv.orgClaimed, not confirmed
Research
The authors put errors in conference papers and show that the systems that check papers are weak against adversarial manipulation.
arxiv.orgClaimed, not confirmed
Research
The method changes only the integer codes of the weights and adds no inference overhead.
arxiv.orgClaimed, not confirmed
Research
The paper reports that the 2 filters cause more personalization failures.
arxiv.orgClaimed, not confirmed
Compute
The authors write that the large model keeps 95 to 99% of its precision when it recomputes only 30% of positions.
arxiv.orgClaimed, not confirmed
Security
Because a deterministic rule was 100% correct on all 6 policies, the model is not necessary.
arxiv.orgClaimed, not confirmed
Benchmarks
Without its guards, the verifier always passes, but the skills that it trains are as good.
arxiv.orgClaimed, not confirmed
Research
The researchers say that less than 1 in 10 of the claims that GPT-6-astra withdraws are overstated.
arxiv.orgClaimed, not confirmed
Benchmarks
3 of the 4 datasets that can target harmful refusal are saturated and cannot show differences between models.
arxiv.orgClaimed, not confirmed
Research
The hosted model follows its policy, but it is sensitive to wording and underconfident.
arxiv.orgClaimed, not confirmed
Research
The authors tell that 2 hosted call sites give the same live result for 28% of the spend.
arxiv.orgClaimed, not confirmed