Purpose
This AI Safety and Monitoring Policy defines the controls implemented to ensure that all AI-generated content published on the Services complies with applicable laws, payment industry requirements, and the Company's internal content standards.
This Policy applies to all AI-powered chat, text generation, and content recommendation features available through the Services.
1. AI Content Monitoring
The Company continuously monitors AI-generated content through a combination of automated moderation systems and human review procedures.
Our monitoring framework includes:
- Real-time moderation of user prompts and AI-generated responses.
- Automated detection of prohibited, illegal, or policy-violating content before it is displayed to users.
- Logging of AI interactions for security, compliance, and quality assurance purposes.
- Risk scoring of conversations based on predefined moderation rules.
- Manual review of flagged conversations by the Compliance and Trust & Safety Team.
These monitoring processes operate continuously and are designed to identify and address policy violations before they impact users, business partners, or payment providers.
2. AI Safety Controls and Prevention Mechanisms
The Services incorporate multiple layers of protection to prevent the generation of prohibited or unsafe content.
These safeguards include:
- Prompt filtering to block requests that violate Company policies.
- Automated moderation of user inputs and AI-generated outputs using AI safety classifiers.
- Restrictions preventing the generation of content involving minors, non-consensual sexual content, exploitation, human trafficking, bestiality, incest, extreme violence, illegal activities, or any other prohibited material.
- Context-aware safety systems that continuously evaluate conversations throughout each session.
- Abuse detection mechanisms to identify malicious or automated misuse.
- Continuous updates to moderation rules and safety filters as new threats or abuse patterns emerge.
Any content identified as violating the Company's policies is automatically blocked before being delivered to the user.
3. Incident Response and Escalation
If prohibited or suspicious AI-generated content is detected, the Company follows a documented incident response procedure.
The response process includes:
- Immediate blocking of the content.
- Automatic logging of the incident.
- Temporary restriction or suspension of abusive user sessions when necessary.
- Investigation by the Compliance and Trust & Safety Team.
- Escalation of high-risk, repeated, or legally significant incidents to Senior Compliance Management.
- Reporting to relevant legal or regulatory authorities where required by applicable law or contractual obligations.
Corrective actions may include:
- Updating moderation rules.
- Improving AI safety filters.
- Restricting or permanently terminating abusive accounts.
- Implementing additional technical safeguards to prevent similar incidents.
All incidents are documented and retained in accordance with the Company's record retention procedures.
4. AI Model Review and Continuous Improvement
The Company performs regular reviews of its AI systems to ensure ongoing compliance with legal, regulatory, and payment industry requirements.
The review process includes:
- Monthly evaluation of moderation performance and detection accuracy.
- Quarterly review of AI safety controls by the Compliance Team.
- Periodic internal testing designed to identify potential safety weaknesses and bypass attempts.
- Analysis of false positives and false negatives to improve moderation quality.
- Updates to prohibited content definitions based on regulatory developments, card scheme requirements, and emerging industry risks.
- Documentation of all material changes made to AI safety controls.
The Company maintains a continuous improvement program to ensure that AI safety measures remain effective against evolving threats.