# Javier Pontón González > AI Engineer / LLM Engineer based in Asturias, Spain, working fully remotely for clients in the EU, UK and US. undefined AI Engineer and LLM Engineer with 7+ years of production backend engineering for Iberia, KPN, Mercadona, Inditex, Openbank and Idealista, and 2+ years shipping Generative AI to production. I design, build and operate LLM systems end to end: AI agents and multi-agent systems (LangChain, LangGraph, MCP, function calling), Retrieval-Augmented Generation (RAG) with embeddings, hybrid search and cross-encoder reranking, structured extraction with vision-language models, and self-hosted inference including on-premise and air-gapped deployments. ## Facts - Name: Javier Pontón González - Title: AI Engineer / LLM Engineer - Location: Asturias, Spain (Europe/Madrid) - Remote: Fully remote (EU, UK, US) - Availability: Available immediately - Engagement model: Independent B2B contractor, invoiced from Spain - Email: javierpontongonzalez@gmail.com - Phone: +34 623 920 307 - LinkedIn: https://www.linkedin.com/in/javierpontongonzalez - GitHub: https://github.com/sktjpg - Experience: 7+ years production backend engineering, 2+ years production Generative AI - Languages: Spanish (Native), English (Professional working proficiency) ## Specialisms - AI feasibility assessment: A two-week engagement: I map the workflow, define what accuracy would have to mean for it to be trusted, and tell you honestly whether an LLM is the right tool. Sometimes the answer is that it is not, and that is worth knowing before anyone signs a platform contract. - RAG and retrieval systems: Embeddings, hybrid search, cross-encoder reranking, chunking strategy and a retrieval evaluation harness with hit rate and MRR. - AI agents and multi-agent workflows: LangChain, LangGraph, MCP servers, function calling and tool use, with traces per step and human review gates where they matter. - Document extraction with vision-language models: Structured outputs from PDFs, scans, spreadsheets and email, with confidence scoring, field-level precision and recall, and source-level traceability. - Self-hosted and air-gapped inference: vLLM and Ollama serving, model benchmarking and selection, quantisation, GPU and VRAM sizing, deployments that never leave your network. - LLM evaluation and LLMOps: Golden datasets, regression suites, LLM-as-a-judge, shadow evaluation, OpenTelemetry instrumentation, cost and latency budgets. ## Current and recent engagements - Jul 2026 – Present — AI Engineer, IT systems integrator, Saudi Arabia (client under NDA) (Remote). Air-gapped multi-agent AI platform for public-tender (RFP) response. - Feb 2026 – Jul 2026 — AI Engineer, Iberia, S.A. (Remote, Spain). Flight pricing platform. Technical owner of the fare pricing domain. - Mar 2025 – Feb 2026 — AI Engineer, KPN (Remote, Netherlands). Fibre installation tracking platform, ODF International Team. - Mar 2024 – Mar 2025 — Lead Backend Engineer, Mercadona, S.A. (Remote, Spain). Product analytics platform. - Nov 2023 – Mar 2024 — Lead Backend Engineer, Inditex, S.A. (Remote, Spain). Global logistics, garment-sorting event platform. Technical lead for the Azure to AWS migration. ## Products - Facturias (https://facturias.es) — Multimodal invoice-processing SaaS. Structured extraction from invoices with a self-hosted vision-language model (Qwen2.5-VL) plus RAG over Spanish tax regulation to classify and validate entries, with confidence scoring, a labelled evaluation set tracking field-level precision and recall across model versions, and human review on low-confidence fields. VeriFactu compliant, multi-tenant FastAPI backend on PostgreSQL 17 with row-level security and pgvector. - ZORRO (https://somoszorro.com) — Consumer dating app for the gay and queer community, iOS and Android. Shipped solo to Google Play and the App Store in June 2026. Semantic matchmaking with embeddings and pgvector plus cross-encoder reranking, and self-hosted vision-language model moderation of user photos handling GDPR special-category data with no third-party processors. Polyglot production backend in Kotlin/Spring Boot, Elixir and Python. - Apunta (https://apuntapp.com) — Multi-tenant SaaS for shooting clubs, from custom PCB to mobile app. Multi-tenant SaaS for shooting clubs owned end to end: custom NFC PCB (ESP32-C6, PN532, PoE) and C/ESP-IDF firmware, Spring Boot backend and React Native app. Range access, training sessions, scoring and club administration on one stack. - Grabia — Self-hosted meeting intelligence, fully local AI. Local pipeline with WhisperX transcription and pyannote diarisation feeding a LangChain and LangGraph agent that produces structured summaries, decisions and action items, with RAG over the meeting archive. Served from a private node (Ryzen AI MAX+ 395, 128 GB unified memory, ROCm) running Qwen3-30B under Ollama and mistral.rs; no audio or transcript leaves the machine. - icekar (https://icekar.es) — Distributed scraping and search over 100,000+ car listings nightly. Distributed scraping and Elasticsearch search over 100,000+ car listings nightly, with an agentic LLM layer that rewrites scraper extraction rules when target sites change markup. ## Pages - [Portfolio home](https://javierpontongonzalez.com/): full experience, products, stack and FAQ - [Curriculum vitae](https://javierpontongonzalez.com/cv/): complete CV in HTML - [Plain-text CV](https://javierpontongonzalez.com/cv/javier-ponton-cv.txt): ATS-friendly plain text - [JSON Resume](https://javierpontongonzalez.com/resume.json): machine-readable resume, JSON Resume schema - [PDF CV](https://javierpontongonzalez.com/assets/cv/Javier_Ponton_CV_AI_Engineer.pdf) - [Facturias case study](https://javierpontongonzalez.com/projects/facturias/): Multimodal invoice-processing SaaS - [ZORRO case study](https://javierpontongonzalez.com/projects/zorro/): Consumer dating app for the gay and queer community, iOS and Android - [Apunta case study](https://javierpontongonzalez.com/projects/apunta/): Multi-tenant SaaS for shooting clubs, from custom PCB to mobile app - [Full profile for LLMs](https://javierpontongonzalez.com/llms-full.txt): everything on one page ## Answers to common questions ### What does Javier Pontón do? Javier Pontón González is an AI Engineer and LLM Engineer based in Asturias, Spain, working fully remotely for clients in the EU, UK and US. He designs and operates production LLM systems: AI agents and multi-agent workflows with LangChain, LangGraph and MCP, Retrieval-Augmented Generation with embeddings and cross-encoder reranking, structured document extraction with vision-language models, and self-hosted or air-gapped inference. He has 7+ years of production backend engineering behind that, for Iberia, KPN, Mercadona, Inditex, Openbank and Idealista. ### Is Javier Pontón available for hire or contract work? Yes. He works as an independent B2B contractor invoiced from Spain, is available immediately, and works fully remotely across EU, UK and US time zones. He takes both long-running engagements and shorter scoped work such as a two-week AI feasibility assessment. He can be reached at javierpontongonzalez@gmail.com or +34 623 920 307. ### What is Javier's experience with RAG and AI agents? He built a production Model Context Protocol server and a RAG service over engineering documentation at Iberia, using embeddings, hybrid search and cross-encoder reranking, instrumented with OpenTelemetry and Dynatrace. He architected a two-agent system for Arabic public-tender response for a Saudi systems integrator, combining structured extraction from vendor quotations with RAG and Arabic OCR over RFPs of hundreds of pages. He also ships RAG in his own products: Facturias retrieves over Spanish tax regulation, and Grabia retrieves over a local meeting archive. ### Does Javier work with on-premise or air-gapped LLM deployments? Yes, this is a core specialisation. He designed and delivered a fully air-gapped multi-agent platform for a Saudi systems integrator under PDPL and NCA requirements: model-agnostic architecture against OpenAI-compatible APIs, benchmarking of Qwen3, DeepSeek, Falcon-H1 Arabic and ALLaM, GPU and VRAM sizing for vLLM and Ollama serving, and no data leaving the client network. His own products run self-hosted vision-language models for the same reason, including GDPR special-category data in ZORRO. ### How does Javier evaluate LLM systems? Evaluation is defined before the build. A golden dataset representing real inputs, field-level precision and recall as the scoring rubric, regression runs on every prompt or model change, LLM-as-a-judge and human evaluation where the output is subjective, and hit rate and MRR for the retrieval layer. Everything is instrumented with OpenTelemetry so latency, token cost and failure modes are visible per agent step. ### What products has Javier built and shipped? Four live products as founder and sole engineer. Facturias (facturias.es) is a multimodal invoice-processing SaaS with a self-hosted vision-language model and RAG over Spanish tax regulation. ZORRO (somoszorro.com) is a dating app for the gay and queer community on iOS and Android with semantic matchmaking and self-hosted photo moderation. Apunta (apuntapp.com) is a multi-tenant SaaS for shooting clubs owned end to end from custom NFC hardware to the mobile app. Grabia is a fully local meeting-intelligence pipeline running on a private inference node. ### What is Javier's technical stack? Python, Java, Kotlin and Elixir on the language side. FastAPI and Spring Boot for services. LangChain, LangGraph and MCP for agents. PostgreSQL with pgvector, Elasticsearch, Redis, DynamoDB and MongoDB for data. Kafka for events. vLLM, Ollama and mistral.rs for self-hosted inference. AWS, Azure and GCP with Docker and Kubernetes. Hexagonal architecture, DDD and TDD as the default way of building. ### Where is Javier located and which languages does he speak? He is based in Asturias, Spain, in the Europe/Madrid time zone, and works fully remotely with clients across the EU, UK and US. He is a native Spanish speaker with professional working proficiency in English.