ExamDumpster

Free NCA-GENL sample questions

Real questions from the NVIDIA-Certified Associate Generative AI LLMs practice bank, with the correct answer and an explanation for each one. No junk, no filler.

Try them in the simulator Same questions, with study, timed and flashcard modes.

Showing 10 of 20 free sample questions.

Question 1Choose one

A data science team is preparing a large text dataset for fine-tuning a Llama 3 model. The dataset consists of 500GB of raw text files. The team needs to perform tokenization and data cleaning as quickly as possible. Which NVIDIA library is specifically designed for GPU-accelerated data manipulation and would be most suitable for this task?

Question 2Choose one

A developer is implementing a Retrieval-Augmented Generation (RAG) system to answer questions about internal company documents. They have already generated embeddings and stored them in a vector database. Which step in the RAG pipeline immediately follows the retrieval of relevant document chunks from the vector database?

Question 3Choose one

An MLOps engineer is deploying a large language model using NVIDIA Triton Inference Server. They observe that under high load, requests with long sequences are causing head-of-line blocking, increasing latency for all subsequent requests. Which Triton feature is specifically designed to mitigate this issue by processing requests out of order?

Question 4Choose 2

A hospital is developing an internal chatbot to help doctors quickly summarize patient histories. To ensure patient privacy and prevent the model from discussing off-topic subjects like celebrity gossip or financial advice, which TWO NVIDIA technologies or techniques should be implemented? (Select TWO)

Question 5Choose one

True or False: Using LoRA (Low-Rank Adaptation) for fine-tuning a large language model involves updating all of the original model's weights.

Question 6Choose one

A research team is fine-tuning a 70-billion parameter model on a single DGX node with 8 GPUs. The full model requires more VRAM than is available on a single GPU. To overcome this, they decide to split the model's layers across the 8 GPUs. What is this distributed training technique called?

Question 7Choose one

When evaluating a text summarization model, a team calculates a score based on the overlap of n-grams between the machine-generated summary and a human-written reference summary. This metric is known as:

Question 8Choose one

A developer is using the NVIDIA NeMo Framework to create a custom conversational AI application. They need to define rules for how the AI should respond to inappropriate user queries and ensure the conversation stays on a specific topic. Which NeMo component is specifically designed for this purpose?

Question 9Choose one

What is the primary function of the self-attention mechanism in the Transformer architecture?

Question 10Choose one

A financial firm is using a generative AI model to create market analysis reports. They are concerned that the model, trained on public data, might inadvertently generate text that is too similar to copyrighted articles, creating a legal risk. Which AI safety problem does this scenario describe?

10 more free samples are waiting

Create a free account to unlock the whole NCA-GENL sample bank, or get full access to all 203 practice questions in the simulator.

Create account