Contents
- how artificial intelligence works and its basic techniques
- how to prepare data for effective machine learning
- when to use classic machine learning models
- application of neural networks and deep architectures
- How to use generative AI and LLM effectively
- AI tools and ecosystem – what it is worth knowing at the start
- AI applications across different industries
- how to assess the quality of AI models – metrics and monitoring
Share
how artificial intelligence works and its basic techniques
Artificial intelligence in practice is mainly based on machine learning, that is, training models on data so that they can predict, classify or make decisions. A model is a mathematical function with parameters (e.g. millions of weights in a network), and training consists of adjusting those parameters so as to minimise the error on the training data. What “tells the model what is good” comes from the loss function, e.g. cross-entropy in classification. The most important goal is not perfection on the training data, but generalisation, that is, good performance on new cases.
The basic ML techniques are divided into supervised learning, unsupervised learning and reinforcement learning (RL). Supervised learning answers the question “how do you predict the label?”, for example whether an email is spam (yes/no) or what the price of a flat will be. Unsupervised learning is about “how do you find structure?”, for example through customer segmentation without predefined categories. RL, in turn, answers the question “how do you choose actions to maximise reward?”, and the model learns from the consequences of decisions instead of from ready-made correct answers.
AI effectiveness depends on whether the model is not overfitted and whether it can transfer to data outside training. When performance on the training set is excellent but clearly worse on the test set, you usually need, for example, regularisation, a larger sample of data or a simpler model. It is also worth distinguishing between an algorithm and a model: an algorithm is a procedure (e.g. gradient descent), while a model is the result of applying the algorithm to data (e.g. a trained logistic regression). Since AI does not guarantee “truth” and can hallucinate or repeat errors from the data, the key is to think in terms of data, metrics and error control, rather than believing in the “magic” of technology.
- 01Model trainingMinimising error on data
- 02Goal: generalisationPerformance on new cases
- 03Types of learningSupervised, Unsupervised, RL
- 04Basic ML techniquesSupervised • Unsupervised • Reinforcement • Predicting labels, discovering patterns, learning through interaction
Summary: The key is to tune parameters for generalisation and choose the right learning technique for the task.
how to prepare data for effective machine learning
Effective machine learning starts with data, because its quality is what most often determines the outcome more than the “trendy” model. In practice, missing values, duplicates and incorrect labels can ruin the metrics even with a well-chosen algorithm. A sensible first step is a data profile: class counts, missing values and outliers, for example in pandas or a tool such as ydata-profiling. If you do not know where to start, start with a data profile and initial cleaning — that is the shortest route to improving model quality.
Proper preparation also includes splitting into train/validation/test so you can reliably check whether the model will cope with new cases. Ratios of 70/15/15 or 80/10/10 are often used, and the test set is left untouched for the final evaluation. Equally important is the definition of the label (target), because a poorly defined target is a common reason why “the model does nothing”. For example, for churn you need a precise definition (e.g. no logins for 30 days), otherwise the model learns inconsistent signals.
Feature preparation answers the question “what should the model see?” and often requires feature engineering tailored to the type of data. For tabular data, typical approaches are categorical encodings (one-hot/target encoding), numerical standardisation and time-based features (e.g. day of the week, seasonality). In imbalanced data (e.g. 1% fraud), accuracy can be misleading, so it is better to rely on precision/recall, PR-AUC and techniques such as class_weight, undersampling or SMOTE. One of the most dangerous mistakes is data leakage, that is, a “leak” of future information into the features, which gives excellent validation and disastrous production.
- Check the quality and structure of the data (missing values, duplicates, outliers; profiling in pandas or ydata-profiling).
- Split into train/validation/test (e.g. 70/15/15 or 80/10/10) and do not “touch” the test set until the very end.
- Define the label so that it matches the real business objective (e.g. a clear churn condition).
- Prepare features: categorical encoding, standardisation, and time features where it makes sense.
- Match the metrics to imbalanced data (precision/recall, PR-AUC) and consider class_weight/undersampling/SMOTE.
- Eliminate information leakage (e.g. features containing future data such as “case closure date”).
- Take privacy into account: data minimisation, masking identifiers (hashing, tokenisation) and compliance with RODO/GDPR for personal data.
when to use classic machine learning models
Classic machine learning models are worth choosing when you are working mainly with tabular data and need quick, scalable results without complex deep learning architectures. Linear regression is well suited to forecasting numerical values (e.g. delivery time, energy consumption) and is both fast and interpretable, although on its own it captures non-linearities less effectively without additional features. Logistic regression is a natural choice for “yes/no” decisions (e.g. fraud vs no fraud) and can return sensible probabilities, provided the data is properly prepared. Another advantage is the ability to explain the impact of features, for example through coefficients or SHAP.
When you want to capture non-linearities without getting into complex mathematics, decision trees and random forests often do the job well, because they capture feature interactions without manually building combinations. In practice, for tabular data, gradient boosting (XGBoost, LightGBM, CatBoost) is often the “gold standard”, because it combines high quality with tools for controlling overfitting. If you need a default choice for prediction in a company or Kaggle-style projects on tables, boosting is usually one of the first candidates. SVM can make sense when there is a clear decision boundary and a moderate number of samples, but on large datasets it can be computationally expensive.
If you want to quickly build a benchmark, K-NN can be a useful baseline (“similar cases have similar outcomes”), although with a large number of records prediction is slow and you usually need standardisation (StandardScaler). When you do not have labels, a classic approach remains clustering (K-means, DBSCAN) for segmentation, and the number of clusters can be chosen using the elbow method, silhouette score and business validation. To “see” data in 2D/3D or denoise features, dimensionality reduction (PCA, UMAP, t-SNE) is used, with PCA more often serving compression and modelling, and UMAP/t-SNE visualising local neighbourhoods. Such a set of methods helps you choose the right tool for the problem without unnecessarily complicating the solution.
- 01Tabular data and scalabilityFast results, simple architectures.
- 02Linear regressionForecasting numerical values.
- 03Logistic regression“Yes/no” decisions, clear explanations.
Ideal for fast, interpretable solutions on structured data without complex deep learning.
application of neural networks and deep architectures
Neural networks and deep learning are worth using when you have a large amount of data and are dealing with complex patterns such as image, sound or text. In practice, deep learning makes particular sense when classic models do not provide good quality without enormous feature engineering. For tabular data, a multilayer perceptron (MLP) is often the baseline architecture, but on small and medium-sized datasets it often loses out to gradient boosting. An MLP can, however, be useful when you have very many records and want to learn representations (embeddings), for example for high-cardinality categories.
When working with images, CNNs are used most often, and they work great for object recognition and can be used, for example, to detect defects on a production line. Standard practices include augmentation (e.g. rotation, cropping) and transfer learning with models such as ResNet or EfficientNet. In sequential tasks (text, time series), RNNs/LSTMs/GRUs dominated for years, but in many applications they have given way to transformers. In forecasting, it is also sometimes the case that specialised models (Temporal Fusion Transformer, N-BEATS) outperform classic LSTMs.
The success of training in deep learning is often determined more by hyperparameters (learning rate, batch size, number of epochs) than by the architecture itself. If training “stalls” or behaves unstably, it is sensible to start by checking the learning rate, data normalisation and stabilisation techniques such as gradient clipping. Overfitting with a small amount of data is most often addressed with regularisation: dropout, weight decay and early stopping (e.g. 5–10 epochs without validation improvement). Computational costs remain a real constraint. A GPU (e.g. NVIDIA RTX 3060/4060) can significantly speed up training, and with large transformers it often makes sense to use the cloud (Google Colab, AWS, Paperspace).
How to use generative AI and LLM effectively
You work with generative AI and LLM most effectively when you treat the model as a context-dependent generator, not as a guaranteed source of truth. An LLM creates text by predicting the next tokens, so it can sound convincing even when it is wrong, because it is optimised for language fluency, not fact-checking. For this reason, it is crucial to ask questions in a way that enforces context, constraints and a verifiable answer format. Prompting usually works best when you provide the role, context, constraints and format (e.g. JSON), and for analytical tasks it helps to ask for assumptions and verification steps.
If you need up-to-date information and answers grounded in specific sources, instead of relying on the model’s “memory” use RAG (Retrieval-Augmented Generation), and only then generation. In RAG, the model first retrieves documents and then generates an answer based on them, which reduces the problem of outdated information and makes it easier to cite sources. For example, a company chatbot can use a vector database (Pinecone, Weaviate, Qdrant) and embeddings (text-embedding-3-large or bge-large). Embeddings, understood as vectors of meaning, also support semantic search and deduplication, because they make it possible to catch similar content despite different wording.
- Set a prompt with the role, context, constraints and answer format. In analysis, ask for assumptions and verification steps.
- When sources and freshness matter, use RAG (document retrieval + generation) rather than relying on generation alone.
- Use fine-tuning mainly for style, formatting, classification or dialogue, rather than as a substitute for factual knowledge without the right data.
- In applications, implement guardrails: content filters, format validation (e.g. JSON Schema), allowlists of tools (tool calling) and red-teaming tests for prompt injection.
Fine-tuning makes sense when you care about a specific style or labelling, whereas it is usually not the best way of “teaching the model new information” without the right data. In such cases, RAG more often wins, and tuning should be treated primarily as a tool for formatting, classification or dialogue. In LLM applications, practical risk reduction is based on guardrails, i.e. content filters, format validation (e.g. JSON Schema), allowlists of tools and red-teaming tests for prompt injection. This set of practices makes it possible to use generative AI for specific tasks while keeping the risk of incorrect or uncontrolled answers in check.
- 01The model is a generatorPrediction, not truth.
- 02Prompt qualityContext, role, format.
- 03Fact-checkingforce assumptions and steps.
- 04RAG supportSpecific sources, freshness.
The key is to guide the model and verify the results, rather than rely on its “memory”.
AI tools and ecosystem – what it is worth knowing at the start
At the start of working in the AI ecosystem, it is most worthwhile to focus on Python and a set of tools for rapid prototyping and repeatable training. A sensible minimum is Python 3.11+, VS Code, Jupyter and an environment manager (conda or venv + pip), because this makes it easier to keep track of dependencies. For classical ML, scikit-learn is usually chosen, and for deep learning PyTorch or TensorFlow/Keras. If you want to build working prototypes as quickly as possible, scikit-learn (Pipeline, GridSearchCV) will usually get you to a result faster than jumping straight into complex networks.
Tools for experiment management solve the down-to-earth problem of “how do I avoid losing track of what worked?” and make it easier to compare models. Solutions such as MLflow, Weights & Biases (W&B) or Neptune record metrics, parameters, models and artefacts, which streamlines auditing and further iterations. For learning and building pipelines, public datasets and models from Hugging Face Hub, Kaggle and OpenML are useful. When you need a dataset “right now”, you can reach for e.g. Titanic/House Prices (Kaggle) or IMDB for text and practise the end-to-end process.
In practice, notebooks (Jupyter/Colab) are great for exploration, but production code is better kept in modules and tests so that it can be maintained without pain. A typical workflow is EDA in a notebook, then refactoring into a package, training in scripts and runs via Makefile/CLI. Containerisation (Docker) solves the “it works on my machine, but not on my colleague’s” problem because it packages versions and dependencies (often CUDA too when using a GPU). When you lack local power, you can use the cloud: Google Colab as an easy start or, in a more production-oriented way, AWS SageMaker, Azure ML and GCP Vertex AI, remembering that GPUs are the most expensive and it is worth setting budget limits and instance auto-stop.
In applications based on LLMs, both model APIs (OpenAI API, Anthropic API, Google Gemini API) and frameworks for building solutions (LangChain, LlamaIndex) come in handy. On the operational side, logging prompts, versioning instructions and regression testing responses are crucial, because even a cosmetic tweak to a prompt can change the application’s behaviour. This approach makes it possible to develop the solution iteratively, without losing control over quality. As a result, the tool ecosystem is not a “list of libraries”, but a set of practices that keep the project under control.
AI applications across different industries
AI genuinely automates processes when a task can be described as classification, prediction, search, recognition or supporting human work in the loop. In customer service, ticket classification, answer suggestions and conversation summarisation are implemented fastest, and a safe starting point is often an assistant mode for the consultant (human-in-the-loop). In marketing, AI supports segmentation, LTV prediction and personalisation, provided there is consistent user data. A practical example is targeting campaigns at the top 5–10% according to a predicted purchase score over 7 days, rather than at everyone.
In finance, typical applications include fraud detection and risk scoring, where interpretability and auditability matter. An explanation requirement often appears (e.g. SHAP/LIME) along with drift monitoring, because model quality can change over time. In industry, computer vision detects defects, counts items and checks label compliance, and an example solution is a camera + YOLOv8 model running in real time. The effectiveness of such deployments depends, among other things, on lighting and data from different shifts, because working conditions directly affect the image.
In HR, AI is sometimes used to search CVs and match skills, but it carries a bias risk, so you should not automatically reject candidates without human oversight and it is worth testing equal treatment across groups. In education, AI works well for tutoring, generating quizzes and personalising materials, and a safer scenario is creating learning plans and exercises with an answer key rather than “writing essays” without verification. In office work, generative AI speeds up the preparation of emails, meeting notes and document analysis, but it requires attention to confidentiality (e.g. enterprise versions such as Microsoft Copilot for M365 or settings that disable training on data). In IT, AI can generate code, tests and documentation (e.g. GitHub Copilot), but you still need a review for security, licensing and compliance with the architecture.
how to assess the quality of AI models – metrics and monitoring
The quality of AI models is assessed by selecting metrics appropriate to the type of task and the cost of errors in a given process. In classification, accuracy, precision, recall, F1 and ROC-AUC are most commonly used, while in regression MAE, RMSE and R² are used. If the question is “which metric is best?”, the answer is simple: the one that most faithfully reflects the impact of false alarms and missed cases. In practice, a metric only makes sense if it genuinely supports a business decision (e.g. setting thresholds or prioritising actions).
The most “tangible” way to understand classification errors remains the confusion matrix, because you can immediately see which types of error dominate. In fraud detection tasks, a high recall is often more important (catching most fraud) even at the expense of lower precision, because a human will still verify some alerts anyway. To ensure the result does not depend on a single random data split, cross-validation (k-fold) is used. For time-series data, rolling validation (time series split) is a better choice, so as not to mix the future with the past.
Robust evaluation starts with a baseline, because it allows you to answer the question of whether the model actually adds value beyond a simple rule. A baseline can be as simple as choosing the majority class, the mean value or a business rule, and if AI does not beat it, you usually need to go back to the data and refine the definition of the problem. Next, robustness tests are carried out, meaning performance is checked under more difficult conditions (noise, missing data, different devices). For example, in computer vision it is worth assessing performance on images with poorer lighting and from a different camera, because that is a common reason for quality drops after deployment.
Post-deployment monitoring involves tracking data drift and prediction quality, because a model may lose effectiveness over time despite good test results. Drift is detected, among other things, by observing feature distributions (e.g. PSI), and in practice an alert is often enough when PSI > 0.2 for key features. When decisions depend on thresholds (e.g. blocking a transaction at p>0.9), probability calibration is also important (Platt scaling, isotonic regression), because “0.9” does not always mean 90% without calibration. To be able to reproduce exactly what was running in production, version the model and data (e.g. MLflow Model Registry and DVC) and record the training configuration.
FAQ
Frequently asked questions
How does artificial intelligence work in practice?
It most often relies on machine learning, where a model learns patterns from data instead of using manually written rules. The goal is generalisation, meaning it performs well on new cases rather than only on training data.
Does data quality matter more than the model itself?
Yes, because gaps, duplicates and incorrect labels can significantly worsen results even with a good algorithm. That is why it is worth starting with a data profile and initial cleaning.
When is it worth using classic machine learning models?
When you are mainly working with tabular data and want to achieve scalable results quickly without elaborate deep learning architectures. In such cases, linear regression, logistic regression, decision trees, random forests and gradient boosting often work well.
What mistakes most often ruin a machine learning model?
One of the most dangerous is data leakage, meaning future information leaking into the features. Overfitting is also a problem, as are poorly chosen metrics, especially with imbalanced data.
When is it better to use neural networks instead of classic models?
When you have lots of data and complex patterns, for example in images, audio or text. Deep learning makes sense when classic models do not provide good quality without substantial feature engineering.
How can you use generative AI and LLMs to limit errors?
The best approach is to treat the model as a context-dependent generator, not as a reliable source of truth. A well-structured prompt, RAG, output format validation, content filters and tests for prompt injection all help.







