   Compiling ring v0.17.14
   Compiling rustls v0.23.40
   Compiling rustls v0.22.4
   Compiling rustls-webpki v0.103.13
   Compiling rustls-webpki v0.102.8
   Compiling tokio-rustls v0.25.0
   Compiling hyper-rustls v0.25.0
   Compiling libsql v0.9.30
   Compiling tokio-rustls v0.26.4
   Compiling rustls-platform-verifier v0.7.0
   Compiling ureq v3.3.0
   Compiling hyper-rustls v0.27.9
   Compiling reqwest v0.13.3
   Compiling autoagents-llm v0.3.7 (https://github.com/kgentic/AutoAgents.git?branch=kremory-combo-2026-06-04#0b9fa9b4)
   Compiling kremory v0.1.6 (/Users/jamessheen/Documents/Projects/Ideas/kremory/crates/kremory)
    Finished `test` profile [unoptimized + debuginfo] target(s) in 11.04s
     Running tests/label_precision_benchmark.rs (target/debug/deps/label_precision_benchmark-37564da61aef374d)

running 1 test
label_precision_benchmark: model=qwen2.5:14b
label_precision_benchmark: Engine::new receives Arc<Ollama> directly (NO ArcChatProvider wrapper). fn model() should return: 'qwen2.5:14b'
test label_precision_gte_0_75_on_mock_interview has been running for over 60 seconds
warn: failed to parse fact JSON after repair: invalid type: null, expected a string at line 1 column 738
  extracted: 'Amazon Robotics', ' (id='amazon robotics') label='Entity'
  extracted: 'Amazon Robotics', 'Boston Consulting Group', 'Northeastern University', 'Morocco', 'France', 'South Korea', 'Samantha', 'Ria', ' (id='amazon robotics boston consulting group northeastern university morocco france south korea samantha ria') label='Entity'
  extracted: 'Amazon Robotics", "entity_type_id": 2}, {' (id='amazon robotics entitytypeid 2') label='Entity'
  extracted: 'Boston Consulting Group', ' (id='boston consulting group') label='Entity'
  extracted: 'Boston Consulting Group", "entity_type_id": 2}]}<tool_call><|im_start|>ῆ Samantha, let's correct the JSON format and ensure there are no duplicates. Here is the properly formatted response for the given text without any duplicate entries or invalid structure issues. Note that each entity appears exactly once with a proper `entity_type_id` assigned as per the provided rules and examples. The entities in this text include ' (id='boston consulting group entitytypeid 2toolcallimstartῆ samantha lets correct the json format and ensure there are no duplicates here is the properly formatted response for the given text without any duplicate entries or invalid structure issues note that each entity appears exactly once with a proper entitytypeid assigned as per the provided rules and examples the entities in this text include') label='Entity'
  extracted: 'Boston", "entity_type_id": 3}, {' (id='boston entitytypeid 3') label='Entity'
  extracted: 'France', ' (id='france') label='Entity'
  extracted: 'Morocco', ' (id='morocco') label='Entity'
  extracted: 'Morocco", "entity_type_id": 3}, {' (id='morocco entitytypeid 3') label='Entity'
  extracted: 'Northeastern University', ' (id='northeastern university') label='Entity'
  extracted: 'Northeastern University", "entity_type_id": 2}, {' (id='northeastern university entitytypeid 2') label='Entity'
  extracted: 'Ria', ' (id='ria') label='Entity'
  extracted: 'Ria", "entity_type_id": 1}, {' (id='ria entitytypeid 1') label='Entity'
  extracted: 'Samantha', ' (id='samantha') label='Entity'
  extracted: 'South Korea', ' (id='south korea') label='Entity'
label_precision: 0.0% (0/10 ground-truth name+label pairs matched)
  15 non-canonical labels (TD-012 indicator)
  non-canonical entities: [("Amazon Robotics', ", "Entity"), ("Amazon Robotics', 'Boston Consulting Group', 'Northeastern University', 'Morocco', 'France', 'South Korea', 'Samantha', 'Ria', ", "Entity"), ("Amazon Robotics\", \"entity_type_id\": 2}, {", "Entity"), ("Boston Consulting Group', ", "Entity"), ("Boston Consulting Group\", \"entity_type_id\": 2}]}<tool_call><|im_start|>ῆ Samantha, let's correct the JSON format and ensure there are no duplicates. Here is the properly formatted response for the given text without any duplicate entries or invalid structure issues. Note that each entity appears exactly once with a proper `entity_type_id` assigned as per the provided rules and examples. The entities in this text include ", "Entity"), ("Boston\", \"entity_type_id\": 3}, {", "Entity"), ("France', ", "Entity"), ("Morocco', ", "Entity"), ("Morocco\", \"entity_type_id\": 3}, {", "Entity"), ("Northeastern University', ", "Entity"), ("Northeastern University\", \"entity_type_id\": 2}, {", "Entity"), ("Ria', ", "Entity"), ("Ria\", \"entity_type_id\": 1}, {", "Entity"), ("Samantha', ", "Entity"), ("South Korea', ", "Entity")]

── PHASE 1 DIAGNOSTIC: Metrics snapshot ──
  rql.entity.label_rejected_total: NO REJECTIONS RECORDED
  Other rql.* counters in snapshot:
    rql.entity_types.default_seed_applied [("group_id", "default")] = 1
    rql.db.insert_episode_with_group_ms [] = 0
    rql.ingest.chunk_count [] = 0
    rql.db.list_entities_count [] = 0
    rql.db.list_entities_ms [] = 0
    rql.extraction.structured_call_attempt [("schema", "EntityListIntegerId"), ("arm", "format_schema"), ("model", "qwen2.5:14b")] = 3
    rql.extraction.structured_call_success [("schema", "EntityListIntegerId"), ("arm", "format_schema")] = 3
    rql.extraction.stage_ms [("stage", "entities")] = 0
    rql.extraction.json_parse_ok [] = 8
    rql.extraction.structured_call_attempt [("schema", "RelTypeList"), ("arm", "format_schema"), ("model", "qwen2.5:14b")] = 3
    rql.extraction.structured_call_success [("schema", "RelTypeList"), ("arm", "format_schema")] = 3
    rql.extraction.stage_ms [("stage", "relations")] = 0
    rql.extraction.structured_call_attempt [("schema", "TripletList"), ("arm", "format_schema"), ("model", "qwen2.5:14b")] = 3
    rql.extraction.structured_call_success [("schema", "TripletList"), ("arm", "format_schema")] = 2
    rql.extraction.stage_ms [("stage", "triplets")] = 0
    rql.extraction.entity_count [] = 0
    rql.extraction.fact_count [] = 0
    rql.extraction.fallback_ladder_step [("schema", "TripletList"), ("from_arm", "format_schema"), ("to_arm", "llm_json_repair"), ("model", "qwen2.5:14b")] = 1
    rql.extraction.structured_call_attempt [("schema", "TripletList"), ("arm", "llm_json_repair"), ("model", "qwen2.5:14b")] = 1
    rql.extraction.structured_call_success [("schema", "TripletList"), ("arm", "llm_json_repair")] = 1
    rql.extraction.json_parse_fail [] = 1
    rql.db.insert_entity_with_group_ms [] = 0
    rql.db.set_entity_embedding_ms [] = 0
    rql.db.insert_episodic_edge_ms [] = 0
    rql.ingest.total_ms [] = 0
    rql.ingest.entity_count [] = 0
    rql.ingest.fact_count [] = 0
    rql.ingest.merge_count [] = 0
    rql.ingest.contradiction_count [] = 0
    rql.db.get_entity_ms [] = 0
── END DIAGNOSTIC ──


thread 'label_precision_gte_0_75_on_mock_interview' (10092212) panicked at crates/kremory/tests/label_precision_benchmark.rs:382:5:
label_precision 0.0% (0/10) is below the 65% R1 threshold.
Pre-Phase-4 (TD-012 broken state) produces ≈ 0%. Post-L1+L2+L3 target >= 65%.
15 entities had non-canonical labels.
Check: StructuredCallBuilder FormatSchema arm is active for model 'qwen2.5:14b'.
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
test label_precision_gte_0_75_on_mock_interview ... FAILED

failures:

failures:
    label_precision_gte_0_75_on_mock_interview

test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 5 filtered out; finished in 268.80s

error: test failed, to rerun pass `-p kremory --test label_precision_benchmark`
