Databricks · DCGAE
Validates the ability to design, develop, and deploy LLM-powered solutions on Databricks, covering RAG application design and data preparation, prompt engineering and retrieval chains, model serving and deployment, evaluation and monitoring for quality and safety, and governance with Unity Catalog.
Practice Questions
620
≈ 13 practice exams
Duration
90 minutes
Passing Score
70%
Difficulty
AssociateLast Updated
Sep 2026
Application Development is the heaviest domain at 30%, followed by Assembling and Deploying Applications at 22%, Design Applications at 14%, Data Preparation at 14%, Evaluation and Monitoring at 12%, and Governance at 8%. This 620-question practice bank is built to match that split, so LLM chain construction, prompt augmentation, and Model Serving deployment scenarios get real depth instead of a token appearance. Watch the freshness of whatever you study from: Databricks refreshed the exam guide on March 18, 2026, and the live exam now covers Agent Bricks, MCP server integration, AI Gateway, Genie Spaces, and ai_query() batch inference alongside the classic RAG material.
Test day means 45 scored multiple-choice and multiple-select questions in 90 minutes, delivered as an online proctored exam through Webassessor. Databricks may mix in unscored pilot questions that are not identified and do not affect your result, with extra time already factored in to cover them. There is no published passing percentage; you receive a pass or fail outcome with a per-domain breakdown, and candidates commonly target roughly 70% to be safe. All ML code on the exam is written in Python, SQL can appear for non-ML data manipulation, and the exam is offered in English, Japanese, Korean, and Brazilian Portuguese.
There are no formal prerequisites, but Databricks recommends six months of hands-on generative AI work plus the Generative AI Engineering with Databricks courses in Databricks Academy. Registration costs $200 USD per attempt, and the credential expires after two years; recertifying means retaking the full exam that is live at that time, which matters here because the content shifts fast. Start with the 30 free questions, then work through the full 620-question bank until your accuracy holds steady across all six domains.
The Databricks Certified Generative AI Engineer Associate certification validates an individual's ability to design, develop, and deploy large language model (LLM)-powered solutions on the Databricks platform. The exam tests practical competency across the full generative AI engineering lifecycle, including decomposing complex requirements into multi-stage reasoning pipelines, selecting appropriate models from both open-source and proprietary ecosystems, and implementing retrieval-augmented generation (RAG) applications using Databricks-native tooling.
Certified professionals are expected to demonstrate hands-on proficiency with key Databricks technologies: Vector Search for semantic similarity and document retrieval, Model Serving for scalable endpoint deployment, MLflow for experiment tracking and lifecycle management, and Unity Catalog for data governance and access control. All machine learning code on the exam is in Python; SQL may appear for non-ML data manipulation tasks. The certification is valid for two years, after which recertification requires retaking the current version of the exam.
This certification is designed for practitioners actively building and deploying AI systems in enterprise environments, including AI Engineers, Generative AI Engineers, LLM Engineers, AI Solution Architects, MLOps Engineers, and Data Scientists with a focus on LLM or RAG workflows. It is particularly well-suited for engineers who work within the Databricks ecosystem and need to demonstrate production-level competency in generative AI solution development.
Candidates should have at least six months of hands-on experience developing generative AI solutions, practical familiarity with Python-based ML pipelines, and working knowledge of frameworks such as LangChain or LangGraph. Experience with Databricks-specific tools—MLflow, Unity Catalog, Vector Search, and Model Serving—is strongly recommended before attempting the exam.
There are no formal prerequisites required to register for this exam; any candidate may attempt it. However, Databricks strongly recommends at least six months of hands-on experience in generative AI solution development before sitting for the certification.
Recommended technical knowledge includes Python proficiency (especially for model pipelines and application orchestration), familiarity with LLM concepts such as context windows, tokenization, and prompt engineering techniques (zero-shot, few-shot, chain-of-thought), and practical experience with LangChain or similar orchestration frameworks. Candidates should also be comfortable using Databricks-native tools including MLflow for experiment tracking, Unity Catalog for governance, Vector Search for embedding-based retrieval, and Model Serving for endpoint deployment.
The exam consists of approximately 45 scored multiple-choice and multiple-select questions to be completed within 90 minutes. It is delivered as a proctored online exam, meaning candidates complete it remotely under live or automated proctoring; no external aids are permitted. The exam is available in English, Japanese, Brazilian Portuguese, and Korean. The registration fee is $200 USD (local taxes may apply).
The passing score is 70%. Databricks notes that exams may include additional unscored items used to gather statistical data for future exam development; these items are not identified and do not affect the final score, meaning the total number of questions delivered may be slightly higher than the 45 scored items. The certification remains valid for two years, after which candidates must retake the current exam version to recertify.
Professionals holding this certification are positioned for roles at the intersection of software engineering and applied AI, including Generative AI Engineer, LLM Engineer, AI Solution Architect, and MLOps Engineer. Generative AI engineering roles command some of the highest compensation in the technology sector, with average salaries reported around $214,000 annually in the United States; the certification directly signals enterprise-grade deployment skills that go beyond prototyping or research experience.
The generative AI applications market is projected to grow at a CAGR exceeding 46% through 2030, and employer demand for engineers who can bridge the gap between experimental LLM work and production-ready Databricks deployments continues to outpace supply. As Databricks is widely adopted across Fortune 500 companies for data and AI workloads, this certification carries strong recognition among employers already invested in the Databricks ecosystem. It complements other Databricks credentials (such as the Data Engineer Associate or ML Professional certifications) for practitioners building a comprehensive Databricks certification portfolio.
5 sample questions with answers and explanations. The full bank has 620 questions, enough for 13 full-length practice exams.
Preview — answers shown1. A legal technology company needs to analyze entire merger and acquisition contracts without chunking. Individual contracts can reach up to 100,000 tokens. Quality is the primary concern and cost is not a constraint. Which Foundation Model best meets this requirement? (Select one!)
Explanation
Gemini 3.1 Pro provides a 1 million token context window, making it well-suited for processing entire contracts of up to 100,000 tokens without any chunking. Llama 3.3 70B supports a 128K token context window, which could be insufficient for the largest contracts and introduces risk of truncation at scale. CodeLlama-34B is specialized for code generation tasks and is not designed for legal document analysis or reasoning over dense prose. BGE-large is an embedding model that produces fixed-dimensional vector representations and cannot generate text or perform analytical reasoning over documents.
2. A Generative AI Engineer at Goldfield Analytics is investigating why some requests are missing from their serving endpoint's Inference Table log during high-traffic periods. Which two fields in the Inference Table schema are most useful for diagnosing this issue? (Select two!)
Multiple correct answersExplanation
The status_code field reveals whether logged requests returned error codes — specifically, requests that produce 401, 403, 429, or 500 status codes may not be captured in the Inference Table at all, which directly explains gaps in the log during high-traffic periods where rate limiting or errors are likely. The sampling_fraction field shows the configured capture rate on a 0–1 scale; a value below 1.0 means only a fraction of requests is being logged intentionally by design, so missing entries may be expected behavior rather than a platform error. The request_metadata field stores custom key-value metadata attached to individual requests and does not indicate why records are absent from the table. The client_request_id and databricks_request_id fields uniquely identify individual requests and are valuable for correlating specific entries, but they do not explain the root cause of records being missing from the log entirely.
3. A Generative AI Engineer at Tailspin Research is configuring MLflow's built-in LLM-as-judge scorer to evaluate retrieved chunk relevance for a regulatory compliance assistant. The judge uses the built-in four-point relevance scale. A retrieved passage explicitly addresses every aspect of the user's compliance query in full. Which score should the judge assign to this passage? (Select one!)
Explanation
MLflow's built-in LLM-as-judge relevance scorer operates on a four-point scale from 0 to 3. A score of 3 denotes a highly relevant passage that explicitly and completely addresses the user's query, which matches the described scenario. A score of 0 is assigned when the passage has no relevance to the query at all. A score of 1 indicates the passage is only partially relevant, touching on the topic surface without addressing the question directly. A score of 2 indicates meaningful relevance with substantive coverage of the query, but stops short of the explicit and complete coverage that warrants the highest rating.
4. A Generative AI Engineer at Prestige Analytics has enabled inference tables on their Model Serving endpoint and wants to analyze captured payload data. They need to identify whether a specific request received an error response rather than a successful generation. Which field in the inference table schema should they examine? (Select one!)
Explanation
The status_code field in the inference table captures the HTTP response code returned for each logged request, making it the correct field for determining whether a request succeeded or encountered an error. The request_metadata field contains additional contextual information about the request source and routing, not the outcome or response code. The execution_time_ms field records how long the endpoint took to process the request, which may be elevated for slow requests but does not directly indicate whether the response was successful or an error. The sampling_fraction field is a configuration value between 0 and 1 that controls what percentage of requests are captured overall; it is a table-level setting, not a per-request outcome field.
5. A content moderation team at Meridian Streaming needs to analyze both short video clips and thumbnail images in a single API call to detect policy violations. The system must handle both modalities natively without preprocessing video into individual frames before sending to the model. Which Foundation Model API model should they select? (Select one!)
Explanation
Gemini 2.5 Flash natively supports all four input modalities — text, image, video, and audio — making it the only option capable of ingesting both video clips and thumbnail images in a single API call without requiring frame extraction or other preprocessing. Claude Sonnet 4 supports text and image inputs within the Foundation Model API but does not support video as an input modality, making it incapable of natively processing video content for policy analysis. Llama 4 Maverick similarly supports text and image inputs only and cannot process video content without a separate frame-extraction step that re-routes the workload outside the model's native capabilities. BGE-large is an embedding model that produces vector representations for retrieval and similarity search; it cannot interpret visual or video content, cannot detect policy violations, and cannot generate structured text output of any kind.
45 scored multiple-choice and multiple-select questions in 90 minutes. Databricks may add unscored pilot questions that are not identified and do not count toward your result; extra time is already built in for them.
Databricks does not publish a fixed passing percentage. You receive a pass or fail result along with a per-domain score breakdown; candidates commonly aim for roughly 70% correct to pass comfortably.
$200 USD per attempt, plus any applicable local taxes. Registration is handled through the Webassessor platform and the exam is delivered online proctored.
Six domains: Application Development (30%), Assembling and Deploying Applications (22%), Design Applications (14%), Data Preparation (14%), Evaluation and Monitoring (12%), and Governance (8%).
No prerequisites are required. Databricks recommends six or more months of hands-on generative AI solution work and the Generative AI Engineering with Databricks self-paced courses in Databricks Academy before attempting it.
Two years. To recertify you must retake the full exam version that is live at that time; there is no shorter renewal exam or continuing-education path.
Yes. The exam guide was refreshed on March 18, 2026 (after an interim February 2026 update) and now tests Agent Bricks, managed and custom MCP servers, AI Gateway, Genie Spaces, custom Scorers, and prompt lifecycle management in addition to the original RAG-focused content.
Databricks Certified Data Analyst Associate
DCDAA · 627 questions
Databricks Certified Data Engineer Associate
DCDEA · 628 questions
Databricks Certified Data Engineer Professional
DCDEP · 628 questions
Databricks Certified Machine Learning Associate
DCMLEA · 630 questions
Databricks Certified Machine Learning Professional
DCMLEP · 622 questions
Databricks Certified Associate Developer for Apache Spark
DCASD · 604 questions
$17.99
One-time access to this exam