Marzieh Derakhshannia scite author profile

Marzieh Derakhshannia

3Publications

6Citation Statements Received

104Citation Statements Given

How they've been cited

How they cite others

104

Affiliations

Montpellier Laboratory of Informatics, Robotics and Microelectronics

Publications

Order By: Most citations

Data Lake Governance: Towards a Systemic and Natural Ecosystem Analogy

Derakhshannia

Gervet²,

Hajj-Hassan³

et al. 2020

Future Internet

View full text Add to dashboard Cite

The realm of big data has brought new venues for knowledge acquisition, but also major challenges including data interoperability and effective management. The great volume of miscellaneous data renders the generation of new knowledge a complex data analysis process. Presently, big data technologies provide multiple solutions and tools towards the semantic analysis of heterogeneous data, including their accessibility and reusability. However, in addition to learning from data, we are faced with the issue of data storage and management in a cost-effective and reliable manner. This is the core topic of this paper. A data lake, inspired by the natural lake, is a centralized data repository that stores all kinds of data in any format and structure. This allows any type of data to be ingested into the data lake without any restriction or normalization. This could lead to a critical problem known as data swamp, which can contain invalid or incoherent data that adds no values for further knowledge acquisition. To deal with the potential avalanche of data, some legislation is required to turn such heterogeneous datasets into manageable data. In this article, we address this problem and propose some solutions concerning innovative methods, derived from a multidisciplinary science perspective to manage data lake. The proposed methods imitate the supply chain management and natural lake principles with an emphasis on the importance of the data life cycle, to implement responsible data governance for the data lake.

show abstract

Life and Death of Data in Data Lakes: Preserving Data Usability and Responsible Governance

Derakhshannia¹,

Gervet²,

Hajj-Hassan³

et al. 2019

View full text Add to dashboard Cite

Mixing Biology and Computer Science Concepts to Design Resilient Data Lakes

Derakhshannia¹,

Laurent²,

Martin

2023

View full text Add to dashboard Cite

Data lakes appeared a few years ago, introduced in particular to meet the challenges of storing and exploiting IoT data. They were first considered as a new technical and commercial tool, sold by the main database software editors. More recently, they have become the subject of research, in particular to define what a data lake should be, what it should provide in terms of services, and how it should be built. In this work, we have tried to return to the origins of data lakes, starting from the name “lake”. We present here how we worked, between biologists and computer scientists, to understand the links between natural and data lakes. In this article, we first explore the links between the disciplines of biology and computer science before declining these links for the particular theme of lakes. This could appear as a work of transferring knowledge from biology to computer science, and a “simple” application of the concepts. However, we had to interact and understand each other’s concepts and issues to align a possible comparison between the disciplines, for example to determine at what scale to establish the biological comparison, from DNA to the more macro system of the animal and plant ecosystem present in a natural lake. For this reason, we are inspired by a hybrid method based on ecological and logistical network topology to propose the resilient structure for the data lake. Thus, we use the Ecological Network Analysis (ENA) as a bio-inspired method and Graph theory as a logistical-inspired framework to study the interdisciplinary resilience strategies for the data lake network.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.