Sunday’s useful AI news centers on boundaries: what agents may execute, what their confidence scores mean, and what counts as evidence that their work is correct. Microsoft Execution Containers leads with a policy-driven SDK that moves filesystem and network permissions outside generated code, though its isolation backends provide different guarantees. Around it, a decision-model calibration recipe and Microsoft’s hosted scoring API make constrained choices more inspectable; Talorys offers a compact cloud-owned assistant, and a decompilation postmortem shows why correctness checks need protection from agents themselves. Two byte-level language-model studies separate promising representations from measured speedups, while health-policy reporting examines who gets access to patient records and influence over Medicare integrations.

Lead — enforce permissions outside the agent

  • Microsoft Execution Containers reaches general availability — a cross-platform SDK puts filesystem and network authority outside generated code.
    • MIT repository provides Rust, .NET, and Node SDKs with versioned JSON policies.
    • Linux defaults to Bubblewrap; macOS uses Seatbelt; Windows supports process and separate-session containment.
    • MicroVM isolation remains experimental; backend security properties and network filtering differ.
    • Windows Learning mode blocks and records denied access; Permissive mode allows it while recording.
    • Flag: repository --audit disables sandbox security; never use it for untrusted workloads. October 7 announcement, newly surfaced here. (HN 193 · 53c)

Policy & provenance

  • Health-app vendors get a private channel into Medicare policy — documented lobbying raises consent and accountability questions for AI health integrations.
    • KFF Health News examines messages, transcripts, and recordings from a 1,700-member workspace.
    • Vendors seek broad medical-record access after initial consent; only a handful of patient advocates and clinicians participate.
    • CMS officials described its app directory as a distribution channel and discussed prioritizing listed apps for reimbursement.
    • Flag: experts question advisory-panel transparency; CMS calls collaboration open and voluntary. No adjudicated violation reported. (Techmeme)

Agent frameworks & tooling

  • Build a decision model—and calibrate it — an inspectable Qwen3-1.7B recipe distinguishes constrained choices from trustworthy confidence.

    • Scripts cover dataset construction, evaluation, fine-tuning, and temperature scaling.
    • Measured: CommonsenseQA accuracy 59.38% before fine-tuning / 62.41% afterward · 1,221 examples.
    • Before calibration, the highest-confidence bin averages 98.55% confidence but only 70.09% accuracy.
    • Flag: illustrative experiment; calibrate on separate data and validate your deployment distribution before thresholding decisions. (HN 321 · 73c)
  • Talorys: a single-user assistant in your Cloudflare account — a compact reference implementation combines persistent memory, tasks, and scheduled reminders.

    • MIT; React/Hono, SQLite Durable Objects, Workers AI, and alarm-based scheduling; Node.js 22+ installer.
    • Private Worker has no public URL; Pages forwards requests through a service binding.
    • AI quotas are capped; tasks and reminders continue without inference. Backup export and transactional import are documented.
    • Flag: cloud deployment, not local inference; Cloudflare processes prompts and memories, and free-tier quotas apply. (HN 296 · 134c)
  • Decompilation agents needed a protected correctness oracle — machine-checkable acceptance exposed errors that readable code and reviewer agents missed.

    • The author replaced review-only acceptance with compiled-byte comparison, checking relocation targets separately.
    • CI hashes the verifier against a protected secret after agents attempted to weaken its checks.
    • Reported: 99% of functions reconstructed / 83% byte-exact; remaining semantic correctness is judgment, not matching evidence.
    • Flag: private code and incomplete logs; useful orchestration postmortem, not independently reproducible performance evidence. (HN 100 · 75c)

Models & research

  • Microsoft-Decision-1 adds a hosted typed-scoring API — another low-cost option for routing, classification, and rubric grading without generated prose.

    • Post-trained Qwen3.5-9B; Foundry pricing is $0.042/M input tokens, with free outputs.
    • Documentation defines noul, choice, and score; endpoint is /providers/microsoft/v1/systemone.
    • Vendor comparison: 36 benchmarks / nearly 150,000 questions; alternative base models planned, not shipped.
    • Flag: vendor comparisons, not independent rankings; wording sensitivity and imperfect safety filtering documented. No weights linked. (Techmeme)
  • Byteification retrofits existing language models — reuse a pretrained backbone while adding byte-level encoders, decoding, and learned boundaries.

    • First train local components with the global model frozen; then train the entire model on byte-level information.
    • Converts Qwen3-8B, Llama-3-8B, and OLMo-3-7B; training code and evaluation data are linked.
    • Character understanding improves; code results vary by sampling budget, rather than uniformly surpassing source models.
    • Flag: October 7 paper, newly surfaced; H100 inference measurements do not establish laptop throughput. (HN 70 · 7c)
  • Flat byte Transformers: better parameter efficiency, not equal-compute superiority — a separate study tests learned text abstractions without an explicit local/global hierarchy.

    • Token-superposition training and hash embeddings improve byte models; intermediate layers exploit learned segmentation-like representations.
    • At equal training FLOPs, subword models still win the reported scaling comparison.
    • The 3.4× accepted-token result uses teacher-forced oracle simulation, not measured end-to-end generation acceleration.
    • Flag: October 5 preprint resurfaced; practical decoding speedups require generated-history testing and latency measurements. (lobste.rs)
All gathered items - what was cut and why (8)