Skip to content
akaturk Academic measurement

OpenAlex topic

Data Quality and Management

This page lists works and academicians tagged with an OpenAlex topic. It is not a YÖKSİS primary or secondary field.

OpenAlex 556 works 23 author topics

Works

556 works

  1. Explainable address matching in online geocoding: filter-based feature selection and ensemble classification 2026

    No abstract yet.

  2. Comprehensive insights into bitemporal databases: a PRISMA-guided systematic literature review 2026

    Abstract Conventional databases support either retroactive or proactive changes, but not both simultaneously, resulting in the loss of valuable historical data. Bitemporal databases resolve this problem by supporting two temporal dimensions simultaneously: valid time and transaction time. This approach enhances data i…

  3. Codebook for demographic variables in the quantitative dataset. 2026

    Codebook for demographic variables in the quantitative dataset.

  4. Predicting the privacy status of potentially private data items using feature selection algorithms 2026

    A privacy-based risk analysis demands a proper identification of whether a data element is private, non-private or potentially private. Though some of the personal characteristics may be inherently sensitive, others acquire sensitivity due to their statistical and semantic association with already known private variab…

  5. Transformer-based feature integration for predictive modeling in multi-source business environments 2026

    Abstract Predictive analytics in multi-source business environments encounter significant challenges when incomplete data sources compromise analytical accuracy. This study presents transformer-based feature integration mechanisms for handling missing data in prediction tasks, validated through a representative case s…

  6. Web Tabanlı Veri Toplama Yöntemleriyle Küresel Firma Bilgilerinin Analizi: İletişim Bilgileri ve Sektörsel Dağılım Üzerine Bir Çalışma 2026

    Günümüz dijital çağında veri, stratejik karar alma süreçlerinin temel bileşenlerinden biri haline gelmiştir. Web scraping yöntemi bu bağlamda, internet üzerinden otomatik veri toplama ve analiz süreçlerinde etkin bir şekilde kullanılmaktadır. Bu çalışmanın temel amacı, küresel ölçekte faaliyet gösteren firmaların ilet…

  7. Improving Fairness in Doubly Imbalanced Datasets 2026

    Fairness has been identified as an important aspect of Machine Learning and Artificial Intelligence solutions for decision making. Recent literature offers a variety of approaches for debiasing, however many of them fall short when the data collection is imbalanced. In this paper, we focus on a particular case, fairne…

  8. Hybrid SE-FAISS: A Novel Semantic Embedding with Facebook Artificial Intelligence Similarity Search-Based Deep Learning Framework for Similar Text Detection 2026

    Efficiently identifying similar textual information is a critical task in modern information technology service desk systems, where ticket volume and redundancy can impede timely responses. In this study, we propose Hybrid SE-FAISS, a deep learning–based framework that combines Semantic Embedding (SE) using a sentence…

  9. A New Multidimensional Data Anonymization Algorithm for Privacy-Preserving Data Publishing 2026

    Developing a privacy-preserving data publishing algorithm that prevents individuals’ identity information from being disclosed without ignoring the utility of the data is still an important goal to be achieved. Finding the optimal balance between data utility and data privacy is an NP-hard problem. In this study, a ne…

  10. SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models 2026

    As Large Language Models (LLMs) for code increasingly utilize massive, often non-permissively licensed datasets, evaluating data contamination through Membership Inference Attacks (MIAs) has become critical. We propose SERSEM (Selective Entropy-Weighted Scoring for Membership Inference), a novel white-box attack frame…

  11. SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models 2026

    As Large Language Models (LLMs) for code increasingly utilize massive, often non-permissively licensed datasets, evaluating data contamination through Membership Inference Attacks (MIAs) has become critical. We propose SERSEM (Selective Entropy-Weighted Scoring for Membership Inference), a novel white-box attack frame…

  12. Bridging Local and Federated Data Normalization in Federated Learning: A Privacy-Preserving Approach 2026

    Data normalization is a crucial preprocessing step for enhancing model performance and training stability. In federated learning (FL), where data remains distributed across multiple parties during collaborative model training, normalization presents unique challenges due to the decentralized and often heterogeneous na…

Academicians

23 academicians