hi there👋, I'm
Pragnyan Ramtha
about me.
- •I'm an AI Engineer specializing in LLM fine-tuning (PEFT/QLoRA), autonomous AI agent systems, and model compression
- •I built Agent7 solo — a no-code platform for long-running AI agents — and previously held #1 on the ARC-AGI public leaderboard
- •I've won 5 hackathons, contributed 70+ merged open-source PRs to projects like openai-node, langgraphjs, and pydantic-ai, and I'm focused on building cost-efficient AI systems that ship to production
- •My core specializations: retrieval-augmented generation (RAG), quantization (GPTQ, mixed-precision), and building autonomous AI agent frameworks., openai-node, langgraphjs, pydantic-ai, promptfoo, haystack, mem0, chainlit, agno, fastmcp, dify-plugins, dstack, opik, mastra
experience.
Founder Remote
at, Agent7
2025 - Present
- Shipped a no-code AI agent platform from zero to production solo, connecting 1,000+ app connectors and orchestrating 20,000+ tools by building the full stack: agent engine, approval system, integrations, and no-code frontend.
- Enabled teams to hand off 20+ hours of weekly repeat work, agents run autonomously for hours to days across research, outreach, CRM, and cross-app workflows by designing a long-running orchestration system with approval-first safety controls.
- Next.js
- TypeScript
- AI Agents
- Full-Stack Development
- SaaS
Software Engineer Remote
at, Learnable India
Apr 2026 - Present
- Built assistive software for blind learners from scratch, shipping a learning portal with screen-reader compatible workflows by engineering accessibility-first interfaces and AI-powered assistance tools.
- Delivered accessible class sessions for visually impaired students, adapting tools and workflows for screen-reader compatibility across real-time sessions by collaborating with educators and TEDx speakers to design inclusive learning experiences.
- Accessibility Engineering
- Accessibility
- Python
- Learning Portals
- AI Assistance
Agentic AI Developer Hyderabad, Telangana - Remote
at, Reputation Dao
Aug 2025 - Jan 2026
- Architected a GCP serverless backend achieving 99.9% uptime by leveraging Cloud Functions and Cloud Run for production-grade AI orchestration.
- Reduced inference latency by 50% across support workflows by engineering a Gemini API response system with optimized prompt caching.
- Boosted accuracy and trust by developing a RAG pipeline utilizing semantic search for real-time documentation retrieval and source attribution.
- Web Development
- Full-Stack Development
- GCP
- Gemini API
- RAG
- Python
ML Engineer (Freelance) Remote
at, Six Axis Studios
Feb 2025 - May 2025
- Researched world-model approaches for architecture workflows to generate CAD-style design outputs from spatial context and architect design intent.
- Prototyped ML pipelines that translated early architectural concepts into structured geometry, creating a faster path from design exploration to CAD handoff.
- Evaluated generated layouts against architectural constraints to improve reliability before model outputs were used in downstream design workflows.
- Artificial Intelligence
- World Models
- CAD Generation
- Machine Learning
- Python
projects.
Agent7
- Built a no-code AI agent platform from scratch, solo. Agents connect to 1,000+ apps and run autonomous sessions for hours or days.
- Handles research, outreach, CRM updates, lead qualification, and cross-app workflows from plain English instructions.
- Next.js
- TypeScript
- AI Agents
- Integrations
- SaaS
AIMO-3: Efficient Reasoning via LLM Fine-Tuning
- AIMO-3 Solver Medalist — fine-tuned Phi-4 (14B) on CoT and TiR datasets to optimize multi-step mathematical reasoning, achieving 90% accuracy on competition benchmarks rivaling 125B parameter models with a fraction of the compute.
- Phi-4
- Fine-tuning
- CoT/TiR
- PEFT
Personality Clone
- Fine-tuned a Large Language model, leveraging PEFT (QLoRA) and contrastive learning on private conversational data to emulate personal response style.
- Implemented a siamese network architecture with cosine similarity loss, which improved semantic embeddings and achieved 92% accuracy in replicating my response style, a 28% improvement over baseline models.
- TensorFlow
- Python
- CUDA
- Transformers
github contributions.
blogs.
How I Built a Custom Harness to Beat the TerminalBench Leaderboard — Using GLMApr 2026
Read full post 13 min read- The harness I built around GLM for TerminalBench, covering entry-point discovery, enforced planning state, trace retention, and parallel subagents.
- Why the leaderboard data convinced me that harness engineering can close the gap between a mid-tier model and frontier agents.
- TerminalBench
- GLM
- Agent Harnesses
- Benchmarking
- Tool Use
Why OpenAI Sent Me $500 for a Research ProjectApr 2026
Read full post 8 min read- How I built XSA4 + EMA + GPTQ-Int6 to place top-5 globally in OpenAI's Parameter Golf challenge.
- Explaining bits-per-byte, Cross-Sparse Attention, EMA smoothing, and what GPTQ actually does under the hood.
- Parameter Golf
- OpenAI
- Cross-Sparse Attention
- GPTQ
- Model Compression
How I Reached #1 on ARC-AGI-2Apr 2026
Read full post 7 min read- How I adapted my parallel agent setup from AIMO3 to hit the top of the ARC Prize 2026 leaderboard.
- I used per-puzzle test-time training, a heavily restricted DFS beam search, and an augmented re-scoring trick to make it work.
- ARC-AGI-2
- Test-Time Training
- Qwen
- Kaggle
- Reasoning
research papers.
Scaling Context Windows to Infinity: A Comprehensive Study of Position Encoding, Attention Mechanisms, Memory-Efficient Inference, and Context Reduction Techniques in Large Language Models 2026
Read paper Academia.edu
A comprehensive analysis of techniques for extending context windows in large language models, examining position encoding strategies, efficient attention mechanisms, and memory-optimized inference approaches to enable processing of arbitrarily long sequences.
Unlocking Societal Trends in Aadhaar Enrolment and Updates: Anomaly Detection and Fraud Risk Prediction 2026
Read paper Academia.edu
A data-driven approach to identify suspicious patterns in India's Aadhaar biometric identification system, utilizing machine learning for anomaly detection and fraud risk prediction in enrollment and update processes.
Speeding Up LLM Inference Using Quantum Computing Techniques 2026
In progress
Exploring quantum-inspired algorithms and quantum computational primitives to accelerate inference in large language models, investigating quantum annealing for attention mechanisms and variational quantum circuits for efficient token generation.
technical skills.
Languages:
Python, TypeScript, Rust, C, SQL, Bash, Go
AI/ML:
PyTorch, Transformers, Unsloth, PEFT/QLoRA, RAG Pipelines, AI Agents, LLM Fine-tuning, CUDA
Full-Stack:
Next.js, React, Node.js, Tailwind CSS, PostgreSQL, REST APIs, SaaS Architecture
Infrastructure:
GCP, Docker, Linux (Arch), Nix, Git, CI/CD
Achievements:
- Top 5, OpenAI Parameter Golf (1.1271 BPB)
- Prev. #1 on ARC-AGI Public Leaderboard
- AIMO-3 Solver Medalist
- 70+ merged open-source PRs
- Winner, NIAT RAG Challenge
- Finalist, Smallest AI × IIT G
- Winner, IEEE Summer of Code 2025
- Winner, Empathy Encryption Hackathon 2025
- Winner, Daydream Hyderabad @ Hackclub 2025
- Top 0.5% Finalist, Shell AI Hackathon 2025
Certifications:
Machine Learning Specialization (Stanford), CS50: Computer Science (Harvard University)
Developer Tools:
Neovim, Arch Linux, uv
Let's work together.
I'm always interested in new opportunities and exciting projects. Whether you have a project in mind or just want to chat about tech, I'd love to hear from you.
Open to AI engineering roles, consulting, and collaboration on ambitious AI projects
Response time: Usually within 24 hours
Pragnyan Ramtha · 2026