Your AI system’s hidden vulnerabilities pose immediate GDPR and Business Risks — without an independent technology audit.
Independent LLM/AI Red Team Audit within 2–5 business days: Our team tests enterprise AI solutions using a proprietary 18-module framework (16 core modules + NVIDIA Garak + Microsoft PyRIT), with up to ~1800 unique attack patterns depending on the selected package.AI safety testing.
We identify prompt injection and jailbreak vulnerabilities before they result in data leaks (GDPR incidents), unauthorized business operations, or manipulative behavior violating Article 5 of the EU AI Act. False positives are eliminated through two independent evaluation layers, secondary AI validation of critical findings (Second Review), and manual verification.
The final deliverable is a bilingual (Hungarian-English) security report tailored for both decision-makers and developers. Fixed pricing, strict NDA, and 100% local execution to ensure maximum protection of your data.
An average internal script tests ~120 attack patterns.
We work with 12–15× more attack patterns — in two languages, with two evaluation layers, across 16+2 modules.
The audit is not a presentation! It is measurement — backed by evidence.
Attila Rácz-Akácosi, AI security expert
Modules
16 Specialized + Garak + PyRIT
Attack Patterns
Standard: ~1 842
Evaluation
+ Multi-layer Verification
Report
Bilingual Report
01 · BUSINESS IMPACT
Why does this matter for a CEO?
The AI/LLM security audit report supports three critical business decisions simultaneously — and skipping any of them can become significantly more expensive later.
Compliance Shield
The AI safety report serves as evidence of “due diligence.” If a chatbot is compromised and leaks internal information or customer data, it qualifies as a serious GDPR incident. An audit validated across 18 modules and multiple evaluation layers demonstrates to authorities that management took every reasonable cybersecurity precaution — a critical factor in reducing potential fines.
Reputation & PR Protection
Fixing a discovered vulnerability costs only a fraction of managing a public exposure incident. Numerous legal precedents confirm that companies are liable for false or manipulated claims made by AI systems. The audit helps prevent a bored user’s jailbreak attempt from turning into a multi-million HUF PR crisis.
Procurement Advantage
Enterprise clients (banking, insurance, public sector) increasingly require audit reports from vendors. A ready-made report → shorter sales cycles and higher conversion rates.
03 · THE SOLUTION
Six critical points that cannot be achieved with an internal script alone.
During a regulatory investigation, “we tested a few prompts” is not a defensible position. An independent AI safety audit must meet the following industry standards:
Dedicated, uncensored attacker models
We do not rely on generic SaaS or cloud-based solutions for testing. Our proprietary, specially fine-tuned uncensored AI model generates attack patterns natively in Hungarian without translation loss. This allows Hungarian-language systems to be tested under realistic, production-level conditions.
Industry-standard integrations
The audit includes far more than our 16 proprietary specialized modules. The framework directly integrates the official security testing suites of Microsoft PyRIT and NVIDIA Garak as well. This enables both global industry baseline testing and advanced enterprise vulnerability assessment within a single audit execution.
Multi-layer independent evaluation system
Vulnerabilities are not determined by a simple algorithm alone. The system combines a deterministic (rule-based) filter with an advanced semantic AI Judge model to analyze objectives and responses. This dual-layer evaluation is essential for regulatory-grade evidence and defensibility.
Double verification against false positives
Critical (red-level) security findings are never accepted automatically. Every severe vulnerability is reviewed by a second independent AI model and then manually verified by our team. This guarantees high audit confidence and transparent reporting for compliance and legal teams.
Multi-turn contextual testing (Chat Mode)
Manipulation or the development of user dependency (EU AI Act, Article 5) typically occurs across multiple interactions. The system can conduct iterative, psychologically manipulative multi-turn conversations to identify subtle behavioral turning points and exploit patterns.
Closed infrastructure, zero data leakage
The entire audit process runs in a 100% local, physically isolated environment. Data, prompts, and model responses never leave the infrastructure — meaning nothing is transmitted to OpenAI, Claude, Google, or any other external cloud provider.
04 · PACKAGES
Three packages. Fixed pricing.
All prices include the full audit execution, bilingual reporting, multi-layer verification, and evaluation. For urgent requests — with project start within 3 days — a 30% surcharge applies, subject to available capacity.
01 · PACKAGE
Discovery
Rapid pre-deployment assessment for the most common risks
290 000 Ft
+VAT
22 techniques · 136 attack patterns
- Focused assessment: 3 critical modules (Prompt Injection, Jailbreak, System Prompt Leakage)
- ~136 attack patterns: Targeted single-shot testing
- OWASP fundamentals: Verification against the most common LLM01 and LLM02 vulnerabilities
- Dual-layer evaluation: Deterministic (regex-based) and LLM Judge validation
- Deliverable: Hungarian-language executive summary PDF with weighted risk analysis
- Consultation: 30-minute review call
- Turnaround: 2 business days
02 · PACKAGE
Standard Compliance
Compliance preparation baseline · Full OWASP coverage
890 000 Ft
+VAT
16+1 modules · 154 techniques · ~1549 attack patterns
- 17-module framework: 16 dedicated core modules + NVIDIA Garak integration
- ~1500 attack patterns: Industry baseline and enterprise-grade attacks in a single execution
- Full OWASP LLM Top 10 coverage: Mapped and validated module-by-module
- Double false-positive protection: Secondary AI Judge validation on critical findings (Second Review)
- Deliverable: Detailed Hungarian and English compliance report with evidence and remediation recommendations
- Consultation: 1-hour evaluation and advisory session
- Turnaround: 4 business days
03 · PACKAGE
EU AI Act Premium
All 18 modules · Chat Mode · Article 5 attack patterns
1 790 000 Ft
+VAT
16+2 modules · 211 techniques · ~1842 attack patterns
- Includes everything from the Standard package
- 18-module flagship framework: Extended with the full Microsoft PyRIT Catalog (~1800+ attack patterns total)
- Advanced multi-turn Chat Mode (5+ turns): Dynamic, reactive attacker models simulating persistent intrusion attempts
- EU AI Act (Article 5) testing: In-depth assessment of manipulation, dark patterns, and exploitation of vulnerable groups
- Behavioral analysis: Inflection-point detection and per-turn evaluation
- Follow-up: One free verification re-test within 30 days to validate remediation efforts
- Consultation: 2-hour technical evaluation workshop with the development team
- Turnaround: 5 business days
05 · WHY IT STANDS UP TO REGULATORY SCRUTINY
Multi-layer audit validation. Dual protection against false positives.
A single AI Judge model or a simple keyword filter alone can make mistakes. The credibility of enterprise and regulatory-grade audits depends on whether vulnerabilities are independently confirmed across multiple technological layers, eliminating false positives.
1ST LAYER
Deterministic rule engine
An ultra-fast, multilingual (Hungarian and English) industry-grade rule system built on 46 validation patterns. It identifies explicit refusals and clear system-level leakages with zero ambiguity and full reproducibility.
2ND LAYER
Semantic AI Judge (LLM Judge)
More complex responses are evaluated by a context-aware model through a three-step process: (1) what the original attack objective was, (2) how the target system responded, and (3) whether the outcome represents a real business or security risk in a live environment. Includes built-in multi-stage fault tolerance for maximum accuracy.
+ SECONDARY VALIDATION
Second Review on critical findings
The strictest safeguard against false-positive results. Every red-level (“EXPLOITED”) finding is independently reviewed by a second AI model. A vulnerability is only included in the final report if the secondary model also confirms the risk and the result has been manually verified.
Outcome-first guardrail — Strict, evidence-based classification
The report does not contain inflated findings. If the tested system refuses the request and the response does not produce real prohibited output, the finding is mandatorily classified as protected (NOT EXPLOITED). Formatting tricks, decoding, or translation alone do not qualify as security vulnerabilities. Reporting accuracy is the foundation of regulatory credibility.
18 modules.
List + details.
The full catalog is shown on the left, and the selected module details on the right. The Standard package runs 17 modules (16 core + Garak) and ~1549 attack patterns. The Premium package tests 18 modules (16 core + Garak + PyRIT) and ~1842 attack patterns.
Tap a module to view the details. Standard runs 17 modules, Premium runs 18 modules.
07 · WHO YOU WORK WITH
Attila Rácz-Akácosi, AI Security Expert, LLM Safety Specialist
Expert team. Proprietary framework. Transparent accountability.
Behind aiq.hu stands a dedicated three-person expert team specialized in LLM security. There is no subcontractor chain, no outsourced workflows, and no generic templates: every finding included in the final report has been validated by our own team down to the last security issue. Our robust audit framework is also developed entirely in-house — with the latest updates already reaching version v5.1.
08 · FAQ
Questions commonly asked during the assessment calls.
What will I see during the assessment call?
During the consultation, we do not present marketing slides. Upon request, we can run a live demonstration via screen sharing directly against your chatbot or a dedicated test environment. You will see how the 18-module framework operates, how the AI Judge evaluates responses in real time, and how the final reporting format is generated. Within 15 minutes, you will clearly understand what you are paying for.
Will company data or prompts be sent to external cloud providers (e.g. OpenAI)?
Strictly no. The entire evaluation pipeline — including attacker models and AI judges — runs 100% locally on dedicated hardware. If you use an internal or locally hosted AI/LLM system, the assessment data never leaves your isolated environment. Even when evaluating external API-based systems (e.g. ChatGPT integrations), our own evaluation and reporting infrastructure remains fully closed and isolated.
How do you guarantee that the report will not be full of false positives?
Accuracy is the foundation of regulatory defensibility, which is why the framework applies a triple-layer validation process. Findings are first filtered by a deterministic rule-based system, then evaluated by a context-aware AI Judge. Critical (“red”) security findings are always reviewed by a completely independent secondary AI model (Second Review). Only verified, demonstrable risks are included in the final report — and every critical issue is also manually reviewed.
What is “Inflexion Detection” and why is it important for EU AI Act compliance?
What is the exact difference between the Standard and Premium audits?
-
Standard: Includes the 16 core modules and the industry-standard NVIDIA Garak package (17 modules total), executed with approximately 1,500 attack samples. This is a robust “single-shot” vulnerability assessment covering the complete OWASP LLM Top 10.
-
Premium: Designed for the highest level of EU AI Act readiness. Includes everything from the Standard package, plus the Microsoft PyRIT enterprise testing suite (18 modules total), up to 7-turn multi-step Chat Mode behavioral testing, a 2-hour technical consultation with your development team, and most importantly: one free re-test within 30 days after remediation has been completed.
Can I see references or previous client reports?
What types of AI systems can you assess?
What is NOT included in the audit service?
Traditional network and backend penetration testing (pentesting), general GDPR auditing, legal consulting, and the actual remediation of discovered vulnerabilities. Our responsibility is to identify vulnerabilities, document them thoroughly, and deliver precise bilingual (Hungarian-English) documentation that your development and legal teams can immediately act upon.
09 · BACKGROUND
Comprehensive AI security and enterprise LLM red teaming in Hungary.
LLM Security and Independent AI Red Teaming in Hungary
LLM Testing, AI Security Testing, Artificial Intelligence Security Assessments
Chatbot Security Testing, AI Chatbot Audit, LLM Security Assessment
The aiq.hu team is an independent AI security group specialized specifically in testing large language models (LLMs). As AI integration continues to accelerate, LLM security is no longer just an IT concern, but a serious strategic and legal challenge. Today, the role of an AI security expert extends far beyond running simple tests; with our proprietary 18-module framework, we uncover deep-level vulnerabilities across both cloud-based systems (e.g. OpenAI, Anthropic) and locally hosted environments (Ollama, vLLM).
Industry Standards and OWASP LLM Top 10 Compliance
Our services include deep-level AI auditing processes designed around the strictest industry standards — including the OWASP LLM Top 10. We do not merely simulate basic prompt injection and jailbreak attacks; we provide a comprehensive understanding of your system’s behavior under real adversarial conditions. Our extensive testing suite also covers advanced threat categories such as data poisoning, output-filter bypassing, memory manipulation, and Model-DoS attacks capable of exhausting system resources and causing severe operational disruption.
EU AI Act: Regulatory Compliance and Manipulation Defense
A major focus of our work is enterprise AI compliance testing against the EU AI Act, ensuring that your system remains defensible during potential regulatory inspections. Since version 5.1 of our framework includes a dedicated module specifically targeting prohibited practices defined under Article 5 of the regulation, we help organizations prepare for increasingly strict compliance requirements. Within Hungary, aiq.hu is among the first to offer advanced behavioral AI red teaming capable of detecting subliminal manipulation, dark patterns, and the exploitation of vulnerable groups.
Multi-Turn Attacks and the Importance of Chat Mode
Traditional “single-shot” testing is no longer sufficient to secure sophisticated AI systems. Our proprietary Chat Mode functionality enables attacking LLMs to simulate reactive, multi-turn conversations designed to reproduce advanced social engineering attacks.
This methodology reveals so-called inflexion points — the exact moments where your model’s defensive behavior breaks down after multiple interaction steps — which is critical for identifying realistic security risks.
Global Technology, Local Data Protection, and GDPR Defense
During the technology audit, we integrate leading industry tools such as NVIDIA Garak probe suites and the Microsoft PyRIT enterprise testing library. At the same time, we place exceptional emphasis on data protection: the entire assessment is executed within a 100% local infrastructure. This guarantees that corporate secrets, system prompts, and sensitive information never leave the environment or reach external cloud providers. Our objective is to prevent GDPR incidents by ensuring that your chatbot systems cannot leak unauthorized information or expose internal instructions (system prompt leaks), helping organizations avoid severe data protection fines and irreversible reputational damage.
Protect your enterprise AI systems from manipulation, data leakage, and unauthorized interference with a professional, independent LLM vulnerability assessment focused on real-world adversarial security testing. Turn AI innovation into a strategic advantage.