Verification layers for agents
Turning agent prose into atomic, typed claims with evidence pointers — then evaluating them with deterministic checks, constraint solvers and SMT solvers — never with another LLM.
Research
We work on the unglamorous, decisive layer of AI: whether an output can be trusted, what a plan will do before it runs, and how a system learns from outcomes that arrive weeks later. Every thread ships into Nikash, Adhva or Swamarg.
Research threads
Turning agent prose into atomic, typed claims with evidence pointers — then evaluating them with deterministic checks, constraint solvers and SMT solvers — never with another LLM.
Joining real-world outcomes that arrive days or weeks later back to the decision that caused them — correlation keys, declarative reward specs, credit assignment and safe shadow → canary rollouts.
Declaring a business as entities, state machines, legal actions, invariants and bitemporal time — so one model yields a validator, a planner, a simulator and a test-case generator.
Business plans fail on time, not logic. Action durations are fitted distributions, exogenous events are processes — and the LLM only proposes candidates that MCTS or beam search scores by simulated rollout.
Small fine-tuned models for high-frequency narrow tasks — HS classification, document typing, entity normalisation — and deterministic code for anything that can be a rule. The burden of proof is on using a model.
Retrieval as a typed query with semantic fallback, not chunk similarity: structured filters, dense and lexical search, explainable rank — and provenance tiers enforced in the schema.
Model strategy
Frontier models for reasoning-heavy steps; mid-tier for bulk extraction; small fine-tuned models for high-frequency narrow tasks; deterministic code for anything that can be a rule.
That discipline is why Swamarg runs its answers, vision moderation, embeddings and voice on our own GPU infrastructure — with no third-party LLM API in the path.
Open problems we are working on
Orchestration runtimes, RAG over documents and “uses the best models” are not differentiators any more. These are.
No standard exists for expressing an industry ontology a runtime can use for retrieval, validation and policy at once.
Nobody has a cross-regime evidence record. Everyone rolls their own audit table.
Golden sets for “did the HS classification hold” compound with every deployment — and whoever owns them owns the category.
A way to declare “this step may be autonomous, this one needs a licensed broker, this one must never be automated.”
Agent identity and scoped tool access remain unsolved in every mainstream framework.
Air-gapped OT, sovereign clouds, no-egress VPCs. Unglamorous — and, in the Gulf, decisive.
Safety by construction
Standards we align with
Threat models and controls are mapped to the frameworks regulators and security teams already use.
Intellectual property
Our founder is lead author on two pending patents in applied AI. Nikash, Adhva, the Swamarg knowledge graph and our self-hosted model pipeline are company-built code and methods.
We collaborate with design partners, researchers and operators who have a real workflow and real outcomes to learn from.