In mass real estate valuation, detecting outliers is crucial for improving model performance. This study examined how different outlier detection techniques impact model performance when using machine learning algorithms. A dataset with 78 variables and 37,269 market samples, enriched with geographic dataset, was created for Tuzla and Pendik districts in Istanbul. Outliers were detected using Boxplot, MAD, and Z-Score techniques. The Boxplot detected the most outliers (3,810), while Z-Score (|Z| > 3) detected the fewest (943). After removing outliers, the dataset was split into training (70%) and test (30%) sets. Then, Random Forest (RF) models were trained, and performance was evaluated using MAE, MAPE, RMSE, and R² metrics. Methods have emerged from different perspectives. Z-Score (|Z| > 3) produced the best results in terms of absolute error and overall performance, achieving the lowest MAE (0.046849) and RMSE (0.065869). The Boxplot method showed consistent performance, with notably low MAPE (12.933) and strong explanatory power (R² = 0.698), making it a good general-purpose technique. MAD provides the highest R² (0.700), making it the most effective in explaining variance. These findings highlight that determining appropriate outlier detection methods can lead to more accurate and robust valuation models by optimizing datasets.
Air pollution remains one of the most critical environmental problems affecting human health, ecosystems, and sustainable urban development, particularly in rapidly urbanizing and industrialized regions. Particulate matter (PM10) and sulfur dioxide (SO2) are among the most critical air pollutants due to their adverse effects on human health, atmospheric processes, and ecosystem integrity. The increasing availability of long-term, high-frequency air quality monitoring data has introduced a large, high-dimensional environmental dataset into environmental studies, necessitating advanced spatial and spatiotemporal analytical approaches to effectively capture pollution dynamics. This study investigated spatiotemporal patterns and temporal shifts in PM10 and SO2 concentrations in the Eastern Marmara Region for 2014, 2020, and 2024. Daily air quality measurements were aggregated into seasonal and annual averages to assess the long-term changes in pollution levels. Geographic Information System (GIS)-based spatial analysis techniques were employed to generate continuous pollution surfaces. In addition, space–time cube (STC) modeling was applied to integrate spatial and temporal dimensions, enabling the visualization of spatiotemporal variability in air pollution levels. The results revealed distinct spatial contrasts and seasonal variations across the study area, with elevated concentrations generally associated with industrial zones and major transportation corridors, whereas lower levels were observed in coastal and less urbanized areas. The results further indicated that the maximum PM10 concentration decreased from 116 µg/m3 in 2014 to 54 µg/m3 in 2020, then rebounded to 91 µg/m3 in 2024. In contrast, SO2 concentrations showed a relatively stable long-term pattern, with annual averages decreasing from approximately 16 to 9 µg/m3 over the same period. Trend analysis revealed a statistically significant decline in PM10, whereas SO2 did not exhibit a significant long-term trend across most districts. The integration of GIS-based methods and space–time cube analysis demonstrates the effectiveness of spatiotemporal analysis of large environmental datasets for understanding complex air-pollution dynamics. It supports informed decision-making for regional air-quality management.
Mass appraisal models are gaining use for improving valuation accuracy, yet their performance remains highly sensitive to how spatial and non-spatial data are structured before training. Clustering algorithms can be used to segment heterogeneous property groups into more homogeneous ones, potentially improving predictive performance. This study investigates the impact of different clustering algorithms, (i.e., K-Means, K-Medians and the Spatially Constrained Multivariate Clustering Algorithm (SCMCA)), on the performance of prominent ensemble learning-based mass appraisal models (i.e., Random Forest (RF), the Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost) and the Light Gradient Boosting Machine (LightGBM)). Using a comprehensive real estate dataset, clustering quality is evaluated using Silhouette, Calinski–Harabasz, and Davies–Bouldin indices, and the performance of cluster-based ensemble mass appraisal models is then compared. The findings indicate that the best performance is achieved with the SCMCA–LightGBM model combination, which reached RMSE = 0.061 and R2 = 0.722. Furthermore, it is determined that clustering-based models provide improvements of up to 7.26% in MAE, 10.61% in MAPE, and 8.40% in RMSE, depending on the combination. The results show that clustering is an effective preprocessing step that can substantially enhance the predictive performance and overall quality of mass appraisal models.
Buildings account for a substantial share of global energy consumption and greenhouse gas emissions, highlighting the need for accurate large-scale building energy assessment. However, in Türkiye, the limited availability of Energy Performance Certificates (EPCs) and detailed 3D architectural building models restricts comprehensive urban energy analyses. This study proposes a novel GIS and Urban Digital Twin (UDT)-based methodology integrating stereo photogrammetry and Machine Learning (ML) to predict annual building energy consumption and evaluate rooftop solar energy potential. Initially, EPCs were integrated with architectural and photogrammetric 3D building models, while building energy parameters were validated using the BEP-TR2 calculation methodology. Validated datasets were then employed to train and compare Random Forest, XGBoost, CatBoost, and Artificial Neural Network models. For buildings lacking detailed architectural models, geometric attributes were extracted from stereo photogrammetric aerial imagery to enable energy consumption prediction. XGBoost achieved the best predictive performance, yielding an R2 of 0.82 using architectural GML data and an R2 of 0.83 using photogrammetrically derived data. Subsequently, rooftop solar radiation potential was estimated and compared with predicted annual energy consumption to assess building self-sufficiency. Results revealed that 27% of buildings were fully self-sufficient, while 20% achieved 50–70% self-sufficiency. The proposed framework demonstrates the potential of integrating UDTs, GIS, stereo photogrammetry, and ML to support scalable urban energy planning, renewable energy integration, and sustainable decision-making.
To support property valuation using three-dimensional (3D) real estate, the present study aims to extend the national spatial data infrastructure (NSDI) with value-affecting factors such as location characteristics, spatial planning, sociocultural information, and building and dwelling characteristics. To this end, the conceptual model proposed in ISO 19152-4 Land Administration Domain Model Part 4: Valuation information is first explored and integrated into the NSDI together with the local characteristics identified as necessary for 3D real estate valuations. The integrated model is then implemented and enriched with georeferenced 3D buildings and floor plans. A case study is then carried out in the city center, where the first 3D models of the building and residential units are produced. In the case study, an overall parametric value score is generated for each residential unit to support real estate valuation in conjunction with GIS analysis and relevant value-influencing factors as determined by a questionnaire completed by valuation experts, academics, and appraisers.
Determining the optimum market sample size for Machine Learning (ML)-based modelling is important to achieve accurate, reliable and economical results in mass appraisal process. Insufficient market samples can lead to the model under-representing the diversity of the market, while many market samples may bring unnecessary computational costs. In this study, a methodological framework and comparison of model performances according to sample size for mass appraisal with ML was designed. A case study was performed in the densely populated Pendik district of Istanbul. 121 variables were determined in structural, spatial and local categories affecting the value of residential real estate, and the datasets were organised in the GIS environment. A total of 142 models were developed with increase of 500 market samples with the Random Forest algorithm, and models were evaluated with performance metrics. The modelling results showed that the models make estimations with high prediction accuracy and increasing the number of market samples improves performance in all metrics. When detecting the model performances based on the market sample size, the general distribution of the breakpoints was quite similar. It was concluded that high-accuracy results can be obtained with fewer market samples when the specific performance range is considered sufficient.
The calculation, management and maintenance of energy performance of buildings (EPBs) are significant in increasing energy efficiency in buildings and reducing greenhouse gas emissions since it is estimated that approximately one third of energy consumption is associated with buildings, and furthermore, three-quarters of the existing building stock is characterized by energy inefficiency. However, in many cases, EPBs are either not calculated or not integrated in a register within national spatial data infrastructure (NSDI). This complicates policy development and planning for both local and national governments, which may result in numerous complications. The objective of this paper is twofold: firstly, to design a building energy data model as an extension of NSDI in Türkiye and then implementing and populating it with real data taken from energy performance certificates from the Tuzla District in Istanbul; and secondly, to develop energy performance prediction models with Machine Learning (ML) algorithms (i.e., Random Forest (RF), Gradient Boosting Machine (GBM), Light Gradient Boosting Machine (LightGBM) and Extreme Gradient Boosting (XGBoost) in order to estimate the overall performance scores of the buildings. The model's findings demonstrated robust predictive accuracy, achieving an R² of 0.818 (XGBoost) and performance metrics of RMSE = 5.153, MAE = 2.886, and MAPE = 3.369. These results substantiate the model's reliability in estimating targeted building energy performance scores. These predictions can be used to provide a comprehensive overview of districts in terms of EPB and inform the development of road maps at the district, city, or national level. Furthermore, the predictions can support the development of EPB-related legislation, facilitate the design of incentive and sanction mechanisms, and promote broader sustainability and climate mitigation goals in a practical manner. Nevertheless, as a limitation of this study, the model has only been tested in a single district, which restricts its generalizability; it should therefore be evaluated in other areas to confirm its applicability.
Son zamanlarda çalışma hayatı ve günlük yaşam biçimlerindeki değişimler daha fazla insanın şehirlere yönelmesine neden olmuştur. Artan nüfus, araç sayısındaki hızlı artışı da paralelinde getirmiş olup şehir içi ulaşım sistemlerini olumsuz etkilemiştir. Bu yüzden sürdürülebilir kentsel ulaşımda büyük öneme sahip otoparkların eksik olması ve uygun olmayan konumlara planlanmasından dolayı problemler meydana gelmektedir. Araçların durağan trafik olarak bilinen otoparklarda zamanının çoğunu geçirdiği göz önüne alındığında, uygun otopark konumların belirlenmesi ile trafik sıkışıklığı ve araçların hareket kabiliyeti optimize edilmektedir. Araç sahipliği ve birim alandaki nüfusun fazla olduğu metropoliten alanlarda ulaşımın sorunsuz sağlanması açısından bu durum büyük öneme sahiptir. Bu çalışmada sürdürülebilir ulaşım planlamasın için otopark uygunluk analizinde Ulaşım, Ekonomi&Finans ve Potansiyel Çekim Özellikleri kriter gruplarında 23 kriter belirlenmiştir. Kriter ağırlıkları için ilgili sektör paydaşlarının katıldığı anket çalışması gerçekleştirilmiştir. Anketler Çok Kriterli Karar Verme (ÇKKV) tekniklerinden En İyi-En Kötü (Best Worst Method-BWM) tekniği ile analiz edilmiş olup, kriter ağırlıkları hesaplanmıştır. İstanbul’un Pendik ve Tuzla ilçeleri ile Kocaeli’nin Gebze, Çayırova ve Darıca ilçeleri çalışma alanı olarak belirlenmiş olup, çalışma alanında kriterlere ilişkin veriler elde edilmiştir. Veriler yakınlık, eğim, bulanık mantık ve ağırlıklı bindirme gibi coğrafi analiz teknikleri ile değerlendirilerek Coğrafi Bilgi Sistemleri (CBS) tabanlı uygunluk analizi gerçekleştirilmiştir. Böylelikle Pendik ilçesinde 5 farklı konumda toplam 18.38 km2, Tuzla ilçesinde 4 farklı konumda toplam 8.55 km2 ve Gebze ilçesinde 6.51 km2 bölgesel uygun alanlar tespit edilmiştir. Sonuçlar kentlerde otoparkların uygun yerlere konumlandırılmasında etkin karar-destek mekanizması olarak değerlendirilebilir. Uygun yerlere planlanmış otoparklar trafik sıkışıklığı ve çevreye salınan karbon emisyonlarının azaltılmasıyla kentsel yaşam kalitesinin artırılmasına katkı sağlayabilir.
Landslides can be considered one of the most severe natural hazards globally, and their management has a key role to inman safety. Landslide susceptibility maps can help the sustainable management of peri-urban areas such as determining the target areas of projects to develop landslide resistant areas, forest planning, infrastructure planning, and land use zoning. This study aims to analyse the relationship between landslide occurrence and its determinants in a peri-urban region, namely Taşlıdere Basin located in Güneysu District of Rize, Turkey, through providing both global and local regression models. To consider spatial non-stationarity, geographically weighted regression was used as a local model. It was found that the local model outperformed the global model in the estimation of landslide susceptibility determinants. The spatial analysis results indicated that almost all variables had heterogeneous variation over the study area. Therefore, this study provides a methodology for understanding the local dynamics of landslide occurrence.
The rapid and uncontrolled development in the urban environment leads to significant problems, negatively affecting the quality of life in many areas. Smart Sustainable City concept has emerged to solve these problems and enhance the quality of life for its citizens. A smart city integrates the physical, digital and social system in order to provide a sustainable and comfortable future by the help of the Information and Communication Technologies (ICT) and Spatial Data Infrastructures (SDI). However, the integrated management of urban data requires the inclusion of ICT enabled SDI that can be applied as a decision support element in different urban problems by giving a comprehensive understanding of city dynamics; an interoperable and integrative conceptual data modelling, essential for smart sustainable cities and successful management of big urban data. The main purpose of this study is to propose an integrated data management approach in accordance with international standards for sustainable management of smart cities. Thematic data model designed within the scope of quality of life, which is one of the main purposes of smart cities, offers an exemplary approach to overcome the problems arising from the inability to manage and analyse big and complex urban data for sustainability. In this aspect, it is aimed to provide a conceptual methodology for successful implementation of smart sustainable city applications within the international and national SDIs with environmental quality of life theme. With this object, firstly, the literature on smart sustainable cities was examined together with the scope of quality and sustainability of urban environment along with all related components. Secondly, the big data and its management was examined within the concept of the urban SDI. In this perspective, new trends and standards related to sensors, internet of things (IoT), real-time data, online services and application programming interfaces (API) were investigated. After, thematic conceptual models for the integrated management of sensor-based data were proposed and a real time Air Quality Index (AQI) dashboard was designed in Istanbul, Türkiye as the thematic case application of proposed models.
The features influencing real estate value in different residential areas and cities are important for spatial economic analysis besides high appraisal accuracy. In this study, a methodology was developed for computer-assisted mass real estate appraisal with a case study implemented through the use of big geographical datasets including 121 features and around 200,000 samples of real estate in Istanbul and Kocaeli (Turkey). Prediction models using the random forest technique were developed for five appraisal zones determined with spatially constrained multivariate clustering. With machine learning and mass appraisal metrics, modelling performance improves in appraisal zones with a lower standard deviation expressing real estate value in neighbourhoods. Since importance levels and ranks of features vary in zones, the mass appraisal should be done with a sufficient number of features.
Elektrikli araçlar, iklim değişikliği ile mücadele, sıfır emisyon hedeflerine ulaşma, sürdürülebilir ulaşım ve enerji verimliliği gibi küresel konulara çevre dostu bir çözüm olarak görülmektedir. Gelişen üretim teknolojileri ve genişleyen satış pazarıyla elektrikli araçların yaygınlaşması beraberinde şarj ekosistemi kavramını getirmiştir. Sunulan bu çalışma, elektrikli araç şarj istasyonlarının uygun yerlere konumlandırılması için coğrafi analiz araçlarının geliştirilmesine odaklanmaktadır. Şarj istasyonları yer seçiminde kritik rol oynayan kentsel parametrelerden çevresel, enerji, kentsel tesisler, sosyoekonomik özellikler ve ulaşım ana kriterler olmak üzere 32 kriter belirlenmiştir. Karmaşık coğrafi verilerin analiz edilmesini kolaylaştırmak ve karar verme sürecinde daha etkili bir yaklaşım benimsemek amacıyla coğrafi bilgi sistemleri ve çok kriterli karar verme teknikleri bir arada kullanılmıştır. Kriterlerin önem sıralaması 20 uzman karar vericiye uygulanan ankete göre yapılmış ve en iyi-en kötü yöntemi kullanılarak ağırlıkları hesaplanmıştır. Pendik ve Tuzla ilçeleri çalışma alanı olarak belirlenmiştir. CBS model geliştirme ortamında 8 adet analitik araç geliştirilmiş, bunlar çeşitli coğrafi veri analiz tekniklerini içermekte, yer seçimi için uygunluk haritalarının üretilmesini sağlamaktadır. Sonuç olarak, farklı çalışma gruplarına uygulanan değerlendirme endeksi sistemi sonuçlarına göre öncelikleri belirlenen kriterler ve geliştirilen analiz araçları, elektrikli araç şarj istasyonu yer seçimi sürecinde karar vericilere ve politika yapıcılara geniş bir bakış açısı sunarak sürecin kapsamlı ve bilimsel bir yaklaşımla yönetilmesini kolaylaştırmaktadır.
Urban land management policies and applications require comprehensive information about real estate development dynamics and value criteria. In this study, geo-analytic tools and a mobile app were produced for residential property valuation in smart cities. Geo-analytical tools were automated to process and use datasets extracted from the databases and 'Smart Real Estate' mobile app using data sets produced by the developed geo-analytics tools shares map services automatically. In this way, users can get real-time location data via web services by triggering geographic analysis tools at certain intervals. It can manage and analyse big geographic data to share smart city services about real estate to citizens.
The rapid increase of population and urban growth have caused huge challenges in ensuring and sustaining environmental and life quality in megacities. In this context, determining, measuring, and analysing the elements that influence city quality of life (QoL) have become important for sustainable urban growth and development. There is a growing interest in QoL assessments, yet reliable and transparent knowledge about the methodology is limited. In this regard, this study aims to contribute to the literature by achieving two distinctive objectives: (1) providing a methodology for investigating the robustness of different weighting approaches to produce a comprehensive and adaptable process for the construction of composite indicators; (2) measuring the urban quality of life at neighbourhood level by including geographical data. Data-dependent statistical methods i.e. Principal Component Analysis (PCA) and Entropy; and multi-criteria decision analysis (MCDA) techniques i.e. Analytic Hierarchy Process (AHP) and Best–Worst Method (BWM) were applied to provide objective and subjective weighting approaches, respectively when calculating the urban quality of life (UQoL) index. The findings showed that there is no considerable difference in the pattern and overall ratio of the weights calculated from different methods; nevertheless, the degree of the weights varies according to the applied method. The results from sensitivity analysis applied to the selected indicator weights covering each method used in the analysis represent the effect of alternative criteria weights on the overall results and the findings point to the weakness of data-dependent methods. The study’s methodology can be applied to similar situations at local, regional, and global scales.
The measurement and monitoring emissions that could be attributed to transportation are highly on demand, since transportation is generally considered as the main source of greenhouse gases. However, several barriers are apparent such as data challenges, national/regional/local emission calculations and lacking efficient tools to measure the performance and progress of counter measures. Furthermore, available open data sources, tools and software are not adequately incorporated. Within this study, in order to facilitate decision-makers on such challenges a web-based geospatial dashboard service is designed to calculate and monitor real-time vehicle emissions. The concepts and the framework is validated in five districts of Istanbul, where Istanbul is prominent with its urban transport challenges. The model, geospatial dashboard, proposed user-friendly and fit-for-purpose, is open-sourced and complies with the national spatial standards. Via utilizing the dashboard, it is possible to monitor emissions from vehicles and uncover spatial patterns with the help of interactive map and graphics. The service provided could help decision makers to perform the technically difficult monitoring process seamlessly, where policymakers could focus on combatting climate change and greenhouse gases. Additionally, the proposed service is easily adaptable to Istanbul and other cities.
The importance of evaluating cities with historical, environmental and liveability parameters have become critical for effective urban management. Modern cadastre and land administration need to register Public-Law Restrictions (PLRs) to ensure the completeness of various transactions. To develop a comprehensive PLR data model for real estate management integrated with National GDI models, this study investigates the urban-related PLRs. An interoperable data model was designed within the Management-Restriction-Regulation Zones data model as an extension with relation to the Cadastre data model. In future cadastral applications, all property-related restrictions can be accessible as in case of the designed data model.
Afet tehlikesi, afetlere neden olan insan ve doğa kaynaklı olaylardır. Afet tehlikeleri ya bir tek olay olarak ortaya çıkar ya da birbirini tetikleyerek peşi sıra gelişir. Afet tehlikeleri birbirini tetiklerse tehlikeler arası ilişkiler karmaşıklaşmakta, zarar görebilirliğin yönü ve boyutu değişmektedir. Tekli afet tehlikelerini bilimsel olarak incelemek oldukça zor iken, çoklu tehlikelerde bu zorluk daha da artmaktadır. Bu çalışma, afetlerde tetikleyen tehlikelerin ve zarar görebilirliğin karmaşık kavramsal yapısını aydınlatabilmek amacıyla gerçekleştirilmiştir. Çalışmada çoklu tehlike ilişkilerinin gösterimi yapılmış; tetikleyen tehlikeleri değerlendirme yöntemlerinden olay ağaçları, etkileşim matrisleri ve olasılıksal modeller tanıtılmıştır. Böylelikle afet risk yönetimi çalışmalarının önemli iki basamağını oluşturan tehlike ve zarar görebilirlik incelemesi tetikleyen tehlikeler kapsamında yapılmıştır.
Following the actions that started in 2003 for the establishment of the National Geographic Information Systems and its infrastructure, the duties of "ensuring coordination between public institutions and organizations, determining the procedures, principles and standards for the production and updating of geographical data and information within the geographical data themes, management, use, access, security, sharing and distribution regarding the National Geographic Information Systems and its infrastructure" were given to the Directorate General of Geographic Information Systems with the Presidential Decrees that came into force in 2018 and 2019."The Data Specification Documents" of 32 data themes, prepared by the thematic working groups in accordance with the national and international standards, were published in the Official Gazette in 2020.It is aimed to provide an interoperable environment by distributing the data harmonized in accordance with the TRGIS standards by institutions and organizations determined as "responsible" and "related" in the National Geographic Data Responsibility Matrix.
Bilişim teknolojilerinin gelişmesiyle, veri üretim teknikleri ve toplanan veri hacmi artmıştır. Akıllı şehir uygulamaları ile sensörler, IoT, internet, giyilebilir teknolojiler gibi farklı veri kaynaklarından akan verilerin yönetimi ve bu verilerden değer yaratmak mümkün hale gelmiştir. Günümüzde toplanan büyük hacimli ve karmaşık verinin yönetimi için geleneksel veri depolama ve yönetim yaklaşımları yetersiz kalmış ve büyük verinin hacim, hız ve çeşitlilik gibi karakteristik özellikleri kapsamında yeni bir yaklaşım ihtiyacı doğmuştur. SQL tabanlı yapısal veri tabanlarının yanı sıra, bu ihtiyaca cevap olarak yapılandırılmamış veriyi yönetmede esnek ve ölçeklenebilir bir çözüm olarak NoSQL veri tabanı sistemleri geliştirilmiştir. Bu çalışmada, akıllı şehirlerde örnek teknolojiler değerlendirilmiş, coğrafi büyük verinin CBS ile entegrasyonu kapsamında hava izleme istasyonlarından elde edilen anlık sensör ölçme verileri kullanılarak NoSQL veri tabanı ortamı olan MongoDB’ de Hava Kalitesi İndeksi (HKİ) hesaplanmıştır. CBS ortamında hava izleme istasyonlarına yakın olan trafik sensörlerinden elde edilen veriler ile ortalama trafik yoğunlukları hesaplanmıştır. Elde edilen sonuçlara göre hava kalitesinin trafik ile ilişkisi belirlenmiştir.
The objective determination of real estate values with current technological approaches has an important role in effective and sustainable real estate management plans. Mass appraisal is the process of valuing a large number of real estate simultaneously instead of evaluating the real estate individually for reducing the loss in terms of time and cost. Machine learning methods, known as advanced estimation approaches, are used to obtain more objective, accurate and fast results in mass valuation processes. In addition, besides sufficient objectivity and accuracy in value determination, these methods can evaluate the relations between the value and the criteria affecting the value holistically. In this context, model successes in mass valuation were examined using Multiple Linear Regression (MLR), Generalized Linear Model (GLM), Support Vector Machines (SVM), Decision Trees (DT) and Random Forest (RF) algorithms. The datasets were divided into 3 groups as Geographic (G), Non-Geographic (NG) and Geographic + Non-Geographic (GNG) and applied separately for modeling with different methods. Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Mean Squared Error (MSE) and R2 were calculated for determining the model measures. Pendik district of Istanbul province was chosen as the application area. By applying different methods with 1475 sampling points representing the real estate sales values for the application area, the performances of the models established were examined with 3 different data sets (G, NG, GNG). Accordingly, RF is the method that gives the highest accuracy while DT and GLM were found as the methods with the lowest accuracy. When the effects of different datasets on the model accuracy were examined, big difference was not observed between the use of the G and the GNG datasets.