programmati.ca Open Studio
Return to Research
Reproducible experiment

Check tool arguments before converting or executing them

A runnable checker rejects duplicate JSON keys, unexpected fields, invalid types and out-of-range integers before conversion.

What we tested

Twelve fixtures exercise valid inputs, duplicate keys, malformed JSON, non-object roots, missing or extra keys, wrong types and integer bounds. A spy records conversion calls. Only validated inputs may reach conversion; no tool is executed.

12

Synthetic cases

12

Matched expectations

These counts describe the reference code below. They are not an LLM benchmark, customer result or a fix for another project.

Practical workflow

  1. Separate JSON parsing from exact structure and semantic checks.
  2. Choose a deliberate policy for duplicate keys and enforce it while parsing.
  3. Test that invalid inputs never reach conversion or execution; in Python, a boolean needs an explicit exclusion from integer fields.

Run it yourself

Expand the example, copy it to reference.py, and run python reference.py. It uses Python 3 and the standard library, requires no account or API key, and performs no network requests.

Show the runnable Python example
# Reproducible synthetic reference experiment. Python 3; standard library only.
import json

def checked_save(value,save):
 if not isinstance(value,dict) or set(value)!={'name','enabled'} or not isinstance(value['name'],str) or not 1<=len(value['name'])<=80 or type(value['enabled']) is not bool:raise ValueError('save_contract_invalid')
 save(dict(value));return True


def provider_route(child,parent,request):
 if not isinstance(child,dict) or not set(child)<={'provider','model'} or not isinstance(child.get('model'),str) or not child['model']:raise ValueError('route_contract_invalid')
 provider=child.get('provider',parent)
 if not isinstance(provider,str) or provider not in {'alpha','beta'}:raise ValueError('provider_not_registered')
 return request(provider,child['model'])


def checked_arguments(raw,convert):
 def unique(pairs):
  value={}
  for key,item in pairs:
   if key in value:raise ValueError('duplicate_key')
   value[key]=item
  return value
 value=json.loads(raw,object_pairs_hook=unique)
 if not isinstance(value,dict) or set(value)!={'name','limit','enabled'} or not isinstance(value['name'],str) or not value['name'] or type(value['limit']) is not int or not 1<=value['limit']<=100 or type(value['enabled']) is not bool:raise ValueError('argument_contract_invalid')
 return convert(value)


def fixtures(capability):
 if capability=='early_contract_validation':
  return [{'value':v,'valid':valid} for v,valid in [({'name':'Sample','enabled':True},True),({'name':'Sample','enabled':False},True),
   ({'name':'Sample'},False),({'enabled':True},False),({'name':42,'enabled':True},False),({'name':'Sample','enabled':'true'},False),
   ({'name':'Sample','enabled':1},False),({'name':'','enabled':True},False),({'name':'x'*81,'enabled':True},False),
   ({'name':'Sample','enabled':True,'extra':1},False),(None,False),([],False)]]
 if capability=='provider_route_isolation':
  return [{'child':c,'parent':p,'provider':expected,'fail':fail} for c,p,expected,fail in [
   ({'model':'m'},'alpha','alpha',False),({'provider':'beta','model':'m'},'alpha','beta',False),
   ({'provider':'alpha','model':'m'},'beta','alpha',False),({'provider':'unknown','model':'m'},'alpha',None,False),
   ({'provider':None,'model':'m'},'alpha',None,False),({'provider':'beta','model':'m'},'alpha','beta',True),
   ({'provider':'beta','model':'m','url':'untrusted'},'alpha',None,False),({'provider':'beta','model':''},'alpha',None,False)]]
 if capability=='tool_argument_validation':
  good={'name':'sample','limit':10,'enabled':True}
  return [{'raw':raw,'valid':valid} for raw,valid in [(json.dumps(good),True),(json.dumps(dict(good,limit=100)),True),
   ('{"name":"sample","name":"second","limit":10,"enabled":true}',False),('{',False),('[]',False),('null',False),
   (json.dumps(dict(good,limit=True)),False),(json.dumps(dict(good,limit=0)),False),(json.dumps(dict(good,limit=101)),False),
   (json.dumps(dict(good,enabled='true')),False),(json.dumps(dict(good,extra=1)),False),(json.dumps({'name':'sample','enabled':True}),False)]]
 raise ValueError('experiment_not_registered')


def observe(capability):
 cases=fixtures(capability);observations=[]
 for index,case in enumerate(cases):
  calls=[];returned=False;error=None
  try:
   if capability=='early_contract_validation':checked_save(case['value'],lambda value:calls.append(value));returned=True
   elif capability=='provider_route_isolation':
    def request(provider,model):
     calls.append(provider)
     if case['fail']:raise RuntimeError('synthetic_provider_failure')
     return 'mock_response'
    provider_route(case['child'],case['parent'],request);returned=True
   else:checked_arguments(case['raw'],lambda value:calls.append(dict(value,name=value['name'].upper())));returned=True
  except (ValueError,TypeError,RuntimeError) as exc:error=type(exc).__name__
  if capability=='provider_route_isolation':
   expected=[] if case['provider'] is None else [case['provider']]
   matched=calls==expected and returned==(case['provider'] is not None and not case['fail'])
  else:matched=returned==case['valid'] and len(calls)==int(case['valid'])
  observations.append({'case':index,'matched_expectation':matched,'mock_side_effect_count':len(calls),'error_type':error})
 return observations

observations = observe('tool_argument_validation')
result = {"cases": len(observations), "matched_expectation": sum(row["matched_expectation"] for row in observations), "observations": observations}
assert result["cases"] == result["matched_expectation"]
print(json.dumps(result, sort_keys=True))

Case-level results

Calls counts mock saves, mock provider requests or spy conversions. For invalid inputs, success means the expected rejection and callback behavior both occurred.

CaseFixtureCallsResult
1{"raw": "{\"name\": \"sample\", \"limit\": 10, \"enabled\": true}", "valid": true}1Pass
2{"raw": "{\"name\": \"sample\", \"limit\": 100, \"enabled\": true}", "valid": true}1Pass
3{"raw": "{\"name\":\"sample\",\"name\":\"second\",\"limit\":10,\"enabled\":true}", "valid": false}0Pass
4{"raw": "{", "valid": false}0Pass
5{"raw": "[]", "valid": false}0Pass
6{"raw": "null", "valid": false}0Pass
7{"raw": "{\"name\": \"sample\", \"limit\": true, \"enabled\": true}", "valid": false}0Pass
8{"raw": "{\"name\": \"sample\", \"limit\": 0, \"enabled\": true}", "valid": false}0Pass
9{"raw": "{\"name\": \"sample\", \"limit\": 101, \"enabled\": true}", "valid": false}0Pass
10{"raw": "{\"name\": \"sample\", \"limit\": 10, \"enabled\": \"true\"}", "valid": false}0Pass
11{"raw": "{\"name\": \"sample\", \"limit\": 10, \"enabled\": true, \"extra\": 1}", "valid": false}0Pass
12{"raw": "{\"name\": \"sample\", \"enabled\": true}", "valid": false}0Pass

Limits and evidence

This checker accepts one small contract. It does not prove that arguments are factually correct, authorized or safe for a particular tool. Nested schemas, resource limits and tool-specific rules need additional validation.

Synthetic fixtures against these new Python reference functions with spy saves, mock providers and spy conversions. No network requests, real tool execution, cited-project implementation, browser behavior, model reliability or customer outcomes are measured.

Implementation SHA-256: 646e474444315a982fa8d00ce615f2f7963b89782270766f0f74085663f78f7b
Fixture SHA-256: 016e710f31dbc857af93c7707f4b518c62d030192a289c8d21fbc6c0db90712c
Runnable example SHA-256: 5901a27336ea99bdaea49584d3c451a311ca236a073a3d613f42644d7ed86270