Skip to content
Shivya Technologies

Research

Research that has to survive production.

We work on the unglamorous, decisive layer of AI: whether an output can be trusted, what a plan will do before it runs, and how a system learns from outcomes that arrive weeks later. Every thread ships into Nikash, Adhva or Swamarg.

2
pending patents in applied AI, lead-authored by our founder
4.5 min → 6–10 s
Swamarg answer time after moving to our own GPU inference
0
third-party LLM APIs in Swamarg's answer path

Research threads

Six threads, one question: can this AI be trusted to act?

R1ships in · Nikash · Gate

Verification layers for agents

Turning agent prose into atomic, typed claims with evidence pointers — then evaluating them with deterministic checks, constraint solvers and SMT solvers — never with another LLM.

Claims & evidenceConstraint & SMT solversSigned attestations
R2ships in · Nikash · Loop

Outcome learning from delayed signals

Joining real-world outcomes that arrive days or weeks later back to the decision that caused them — correlation keys, declarative reward specs, credit assignment and safe shadow → canary rollouts.

Credit assignmentReward specsPolicy rollout
R3ships in · Adhva

Executable world models

Declaring a business as entities, state machines, legal actions, invariants and bitemporal time — so one model yields a validator, a planner, a simulator and a test-case generator.

State machinesInvariantsBitemporal clock
R4ships in · Adhva · Planner

Planning under stochastic time

Business plans fail on time, not logic. Action durations are fitted distributions, exogenous events are processes — and the LLM only proposes candidates that MCTS or beam search scores by simulated rollout.

Monte Carlo rolloutsMCTS / beam searchFitted durations
R5ships in · Platform · Swamarg

Domain-tuned small models

Small fine-tuned models for high-frequency narrow tasks — HS classification, document typing, entity normalisation — and deterministic code for anything that can be a rule. The burden of proof is on using a model.

LoRA / QLoRASLMsSelf-hosted serving
R6ships in · Swamarg

Verified knowledge graphs & hybrid retrieval

Retrieval as a typed query with semantic fallback, not chunk similarity: structured filters, dense and lexical search, explainable rank — and provenance tiers enforced in the schema.

Knowledge graphHybrid searchProvenance tiers

Model strategy

The burden of proof is on using a model.

Frontier models for reasoning-heavy steps; mid-tier for bulk extraction; small fine-tuned models for high-frequency narrow tasks; deterministic code for anything that can be a rule.

That discipline is why Swamarg runs its answers, vision moderation, embeddings and voice on our own GPU infrastructure — with no third-party LLM API in the path.

  1. 1Deterministic codeAnything that can be a rule
  2. 2Small fine-tuned modelsHigh-frequency narrow tasks
  3. 3Mid-tier modelsBulk extraction & classification
  4. 4Frontier modelsReasoning-heavy steps only
  5. prefer ↑ · escalate ↓ only when the task demands it

Open problems we are working on

Where the field is unsettled — and where we intend to lead.

Orchestration runtimes, RAG over documents and “uses the best models” are not differentiators any more. These are.

Open · 01

Domain semantic layers

No standard exists for expressing an industry ontology a runtime can use for retrieval, validation and policy at once.

Open · 02

Portable decision provenance

Nobody has a cross-regime evidence record. Everyone rolls their own audit table.

Open · 03

Domain evaluation datasets

Golden sets for “did the HS classification hold” compound with every deployment — and whoever owns them owns the category.

Open · 04

Constrained-autonomy semantics

A way to declare “this step may be autonomous, this one needs a licensed broker, this one must never be automated.”

Open · 05

Least-privilege agent identity

Agent identity and scoped tool access remain unsolved in every mainstream framework.

Open · 06

Hostile deployment environments

Air-gapped OT, sovereign clouds, no-egress VPCs. Unglamorous — and, in the Gulf, decisive.

Safety by construction

Rules the platform will not break.

  • A licensed gate can never be bypassed.
  • Restricted-party screening hits are never auto-cleared.
  • No writes to OT control systems, ever.
  • Data residency is enforced as policy, not convention.
  • Undeclared tools are denied by default.
  • Every deployment ships with a kill switch.

Standards we align with

Built to be inspected.

Threat models and controls are mapped to the frameworks regulators and security teams already use.

OWASP Agentic AI Security Top 10NIST AI RMF + GenAI ProfileISO/IEC 42001MITRE ATLASSingapore Model AI Governance Framework for Agentic AIOpen GenAI tracing conventions

Intellectual property

Our founder is lead author on two pending patents in applied AI. Nikash, Adhva, the Swamarg knowledge graph and our self-hosted model pipeline are company-built code and methods.

Working on a hard verification or simulation problem?

We collaborate with design partners, researchers and operators who have a real workflow and real outcomes to learn from.

Talk to our research team