← Back to blog
Tutorial

Custom models: Finetuning

00 Dataset Preparation V2

Finetuning models

Stop creating by coincidence—start influencing your AI-generated images.

A custom-trained AI model allows you to generate images with intention rather than relying on chance. Instead of pulling a lever on a slot machine and hoping for a useful result, you gain control over the creative process—guiding the AI to produce visuals that align with a specific style, concept, object, or person.

 

A Cookbook Metaphor: Understanding AI Model Training

 

Think of training an AI model as writing a cookbook dedicated to making pizza. Each page contains:

• An image of a pizza (representing an image in the dataset)

• A list of ingredients and instructions (representing the caption associated with the image)

Once the cookbook is complete, you can use it to create unlimited variations—but the result will always be a pizza. Similarly, a well-prepared dataset ensures that the AI generates images that adhere to your intended theme.

 

Different Types of training

we provide 4 types of settings that should produce good results by default. 

• Style (e.g., cyberpunk, watercolor painting)

• Concept (e.g., futuristic cities, surreal dreamscapes)

• Object (e.g., electric cars, medieval swords)

• Person (e.g., a historical figure, a fictional character)

Choose the one that fits your model and the settings will be adjusted automatically. For further information about the dataset preparation keep reading. 

Uploading Images

For effective training and high-quality results, it is crucial to build a well-structured dataset. In the context of AI training, a dataset consists of a curated collection of images that define a particular visual characteristic.

01 02 03

Checklist for a High-Quality Dataset

To optimize training outcomes, consider the following best practices:

• High-resolution images (minimum 512x512, ideally 1024x1024) ensure better detail retention.

• Diversity and consistency: Images should be varied in content but consistent in artistic style or subject.

• Quantity vs. quality: More images do not always lead to better results. For LoRA training, 20–40 images typically suffice for most styles or concepts, whereas larger datasets (50+ images) may be needed for broader adaptability.

• Variation in composition: Ensure a mix of angles, poses, backgrounds, and lighting conditions to improve generalization.

• Style LoRA training: If training a style-based model, ensure all images share the same artistic style while depicting different subjects.

 

Cropping - setting the right focus

A well-prepared dataset requires focus. When training a custom AI model, the system must understand what the dataset is truly about. AI learns best through repetitiveness and quality, gradually recognizing the concept you want it to internalize. Cropping is a powerful technique to refine this focus.

By strategically cropping images, you can emphasize the subject of interest, ensuring that the AI model consistently recognizes the key visual elements. This is particularly useful in cases where:

• The object in focus is surrounded by distractions or unnecessary background elements.

• The subject is too small within the frame, making it harder for the AI to learn key details.

• You want to establish uniform framing across multiple dataset images to reinforce specific proportions and compositions.

A well-cropped dataset improves the clarity and precision of AI-generated images by reducing unnecessary noise and guiding the AI’s learning process more effectively. When combined with high-resolution images and consistent framing, cropping can help ensure that your dataset remains structured and optimized for training.

 

What is Weighting?

Weighting is a technique that allows users to control the influence of specific parts of a dataset during AI model training. By assigning different weights to images or concepts, users can ensure that certain elements receive more or less emphasis in the final model. This technique is particularly useful when training multiple concepts within the same run or balancing different image compositions.

Weighting can be applied with or without trigger words. In some cases, each bucket (or category) of images may have a unique trigger word, while in others, the same trigger word can be used across multiple buckets. In our interface a bucket is a folder.

First things first: What is a trigger word?

A trigger word is embedded in the dataset during training. It consists of one or more words that signal the system: “If I prompt XY, then XY should occur.” Typically, it is a unique sequence of letters or numbers that do not appear in natural language. Why? Because if a commonly used word serves as a trigger, the base model may associate it with an existing concept, increasing the risk of unintended results. By choosing your own trigger words, you gain control over which word will activate a specific behavior when generating an image.

What is Weighting Used For?

Weighting is particularly useful in two common scenarios:

Case A: Training Multiple Concepts in One Run

Users who want to train different but related concepts within a single training session can use weighting to balance their impact. For example:

• Folder 1: 20 images of a winter scenery in Illustration Style X

• Trigger word: “wnterscn”

• Folder 2: 20 images of a summer scenery in Illustration Style X

• Trigger word: “smmrscn”

• Folder 3: 20 images of an autumn scenery in Illustration Style X

• Trigger word: “atmnscn”

• Folder 4: 20 images of a spring scenery in Illustration Style X

• Trigger word: “sprngscn”

Illustration Balanced

Since each bucket contains an equal number of images, they will have a similar influence on the training process. However, if one bucket had fewer images, weighting could be adjusted to balance its contribution.

Case B: Training a Model with Both Detail Shots and Full Shots

Another use case is when a user wants to train both full-body images and close-up details while maintaining a proper balance. For example:

• Folder 1: 30 images of a man in different outfits and environments

• Trigger word: “mnmdlflux”

• Folder 2: 10 close-up images of the same person focusing on skin texture

• Trigger word: same as in folder 1

If both buckets were given equal weight, the close-up details might be underrepresented. Adjusting the weighting allows the second bucket to have a more significant impact on the model’s ability to generate accurate skin textures.

What Do the Weighting Numbers Represent?

The numbers assigned in weighting correspond to repeats, meaning how often an image set is included in training. 

Understanding Repeats

• 1 dot = 1 repeat

• 5 dots = 5 repeats (higher influence)

Illustration Inbalanced

Applying Weighting in Training

In Case A, where all buckets initially have the same number of images:

• The weighting could be evenly set to 1, ensuring equal influence.

• However, if the Spring scenery bucket only had 10 images instead of 20, the weights could be adjusted:

• Spring folder → 2

• Other folders → 1

• This adjustment ensures that the underrepresented concept maintains a similar influence in training.

In Case B, where a user wants a well-balanced personal model:

• Bucket 1 (full-body images) → weighted at 2

• Bucket 2 (close-ups) → weighted at 3

This means that while full-body shots still have more overall influence, close-up details are slightly increased in importance to improve facial or skin texture accuracy.

 

Captioning: The Key to Precision

Each image in the dataset is paired with a short textual description, known as a caption. This caption describes what is displayed in the image, allowing the AI to associate visual patterns with linguistic concepts.

When the system analyzes an image, it also processes its caption to understand the training context. This enables more precise prompt-based image generation. For example, after training a model with a dataset labeled accordingly, prompting “a three-quarter rear shot of a (Triggerword of the model) standing in the middle of the jungle” will yield results based on the dataset’s learned attributes.

While captioning is not strictly necessary, it significantly enhances control over the generated images. The choice between captioned vs. captionless datasets determines the degree of precision in AI-generated outputs.

Captioned Dataset

• Enhances prompt responsiveness

• Improves contextual adaptation, allowing the model to perform well across different contexts

• Strengthens prompt adherence, ensuring nuanced descriptive elements influence results

Captionless Dataset

• Focuses purely on visual patterns without associating them with specific textual descriptions

• Makes prompt-based control more challenging, requiring greater reliance on base model interpretation

• Results in reduced precision in guiding output through text prompts

Thus, the choice depends on whether you prioritize versatile control (captioned dataset) or visual reproduction within the dataset’s context (captionless dataset). 

How to Write Effective Captions

A simple but effective rule:

• Describe in detail what you don’t want the model to learn

• Keep descriptions general for what you do want the model to learn

This may seem counterintuitive, but let’s examine an example:

04

Consider a dataset featuring a boy in different environments:

• Image 1: The boy in a desert

• Image 2: The same boy in a dark forest

By captioning both images as “a boy”, the model learns that “the boy” always has dark hair, large eyes, a green glow on his chest, combat boots, and a beige jacket. However, since the backgrounds differ, the model understands that “the surroundings” can be variable.

To further refine control:

• If you want the boy to later wear different types of clothing, use “a boy in rugged clothing.”

• If you want a car to later appear in multiple colors, explicitly state the car’s color in the training dataset.

 

Captioning Structure Example

A good caption follows this format:

(Subject) (Action) (Description of Surroundings)

Example:

“A boy standing in a dark forest, dark blue hue, illustration style”

For even stronger prompt control, incorporate the trigger word assigned to the trained model:

“A tlrflx boy standing in a dark forest, dark blue hue, illustration style”

 

Training Settings in Custom AI Models

Training a custom AI model can be overwhelming. With countless settings to tweak and numerous tutorials offering conflicting advice, it’s easy to feel lost. The core issue? Most guides focus on specific use cases, making their recommended settings unsuitable for broader applications.

To simplify the process, we suggest beginning with a set of default training settings. These provide a stable foundation for achieving a reasonable baseline performance. From there, you can gradually refine adjustments to match your unique needs.

What Happens When Training Settings Are Changed?

Modifying training parameters can have significant effects on model performance. Understanding potential outcomes can help you make informed adjustments.

1. Overtraining

• AI models have an optimal learning point, often referred to as coherence—the moment when the model has grasped a concept fully.

• Training beyond this point can lead to artifacts and degraded performance, where the model becomes too rigid or memorizes noise instead of learning patterns.

2. Undertraining

• If a model is not trained long enough or lacks sufficient data, it fails to generalize concepts and cannot produce meaningful results.

• Undertraining often results in incomplete or inaccurate outputs.

3. Errors & Instabilities

• Machine learning models are complex systems. Adjusting settings incorrectly can lead to unexpected errors.

• Parameters such as batch size, learning rate, and optimizer choice influence stability—setting them too high or low can cause failures.

Beginner Tip: Stick to Defaults

For those new to model training, we strongly recommend focusing on data quality rather than tweaking advanced settings. A well-curated dataset is often more impactful than minor hyperparameter adjustments.

 

Sample Prompts in AI Model Training

During the AI model training process, sample prompts are essential for tracking and evaluating the model’s development. These prompts generate images at different training stages, providing users with a preview of how the model evolves over time.

Why Use Sample Prompts?

1. Training Progress Overview: Sample prompts allow you to visualize different stages of the model’s training, making it easier to assess improvements and potential weaknesses.

2. Result Comparison: By examining generated images, you can compare outputs from various training checkpoints and select the best model version.

3. Trigger Word Verification: Ensuring that your trigger word appears in generated outputs is crucial for model accuracy and consistency.

Auto Sample Prompts – A Convenient Option

If you’re unsure what to write, simply press “Auto Sample Prompt” to generate prompts tailored to your training setup. This ensures that you always have structured inputs to monitor your model’s learning process.

By utilizing sample prompts effectively, you gain a structured way to evaluate and refine your AI model, ensuring high-quality, predictable results.

Conclusion

By thoughtfully preparing your dataset, you can significantly increase the control and accuracy of your AI model’s outputs. Whether you’re aiming for stylistic consistency, object accuracy, or character replication, a well-crafted dataset ensures that your AI generates images aligned with your creative vision.

Happy training!